Skip to content

Metadata is produced, not reconstructed

Data archaeology

A catalogue project starts. A small team interviews data owners, reads pipeline code and surveys running systems to work out who owns each dataset, what it means and where it comes from.

By the time an entry is written, the schema has changed twice and the owner has moved teams.

Always stale, always manual

  • Metadata gathered after the fact is expensive to collect and goes stale quickly.
  • People retype facts that are already known somewhere else.
  • The catalogue becomes one more system to maintain by hand, and one more place that disagrees with reality.
  • Discovery suffers, because consumers stop trusting what the catalogue says.

Metadata as exhaust, not archaeology

Metadata should originate from the product lifecycle (declaring, building, deploying and operating the product) rather than depend on retrospective curation. People add only what genuinely needs their judgement.

A catalogue inferring ownership, meaning and lineage long after the fact is repairing missing engineering discipline.

The picture

Metadata as archaeology versus metadata as exhaust Top, archaeology: data is produced, months pass, a curator interviews owners and mines code, and a catalogue entry is written that is already going stale. Bottom, exhaust: declare, build, deploy and operate each emit metadata as a by-product, flowing continuously into discovery. A person adds only meaning and judgement calls. Archaeology · reconstructed afterwards Data produced months pass Curator interviews,mines code Catalogue entryalready stale Exhaust · produced by the lifecycle Declare Build Deploy Operate owner, domain schema, contract lineage, version quality Discovery and catalogue a person adds only meaning and judgement calls
Above: metadata reconstructed after the fact. Below: metadata emitted at every step of the lifecycle.

Declaring a new product

Illustrative example

A team declares a new returns data product. Its owner, domain, default classification and retention rule are pre-filled from what is already known about the team and the source system. The person declaring it writes only what needs judgement: what a return means, and whether the suggested classification is right.

Building it records the schema and contract version. Deploying it records lineage. Running it publishes quality results. Nobody fills in a catalogue form, and the discovery page is never older than the last deployment.

In detail

Scaffolded, then completed by people

Metadata can be pre-populated from domain, product, platform and policy context at the moment a product is declared, then completed by a person only where judgement is required: meaning, ownership decisions, classification calls.

Discovery projects; it does not own

A catalogue or discovery experience should present the evidence the lifecycle produces, not become a separately maintained source of truth.

Not this

  • Not "automate every field". Some metadata depends on human judgement.
  • Not a licence to fake completeness. Scaffolding must not manufacture a false appearance of a finished record.

Unresolved

  • What happens to the existing estate that was never declared this way? How much is worth reconstructing?

Join the argument

Agree, challenge or add evidence

Architecture improves when its assumptions are challenged. Say which kind of contribution you're making:

  • Challenge I disagree because…
  • Evidence We've seen this too…
  • Question How would this work when…?
  • Alternative Another way to approach this…
  • Extension This also implies…

Please keep clients, employers and colleagues unidentifiable. Strong contributions may be quoted and credited on this site. Read the contribution guidelines and the privacy notice.

Comments are provided by giscus and stored on GitHub. Posting requires a GitHub account.

Chapter 3 begins · Data as products with promises Data product What turns a table someone can query into something another team can depend on? Quality starts at the producer