Skip to content

Put computation where it belongs

Repairing grain in every report

Sales land one row per order line. Managers ask for revenue per customer per month, and for the number of distinct customers. Every report joins, removes duplicates and counts distinct customers on the fly, each in a slightly different way.

The mismatch between the grain data arrives at and the grain questions are asked at is real. The problem is that it is repaired again inside each report.

The same repair, paid for many times

  • The same computation is built, paid for and maintained in every report, and each copy can drift.
  • Each new consumer triggers another bespoke copy or pipeline instead of reuse. The estate grows into a ball of mud one consumer at a time.
  • A slow, high-cardinality distinct count over data at the wrong grain gets treated as a performance-tuning problem, when it is a modelling signal.
  • When an AI assistant is added, it is asked to repeat all of this reasoning by itself.

Work flows left

Recurring analytical computation (grain, deduplication, joins, classification and shared business rules) belongs in the data product's serving model and the query layer beneath it. It should never be re-implemented in each consumption technology.

Three responsibilities stay separate:

  • Semantic authority decides what something means.
  • The governed analytical and query capability decides how it is correctly evaluated.
  • The consumption technology decides how the result is requested, used and presented. It is deliberately the smallest of the three.

The picture

Where analytical work belongs Five layers from left to right: source domain (business events, ownership, meaning); data product serving model (right grain, shared rules, additive facts); query and access layer (filter, join, aggregate, push down); presentation tool (navigation and display logic); consumer, such as a BI tool, API, notebook or AI assistant (ask, use, present). An arrow beneath points left: push work as far left as it can be done correctly and reusably. 01 Source domain Business events Ownership Meaning 02 Serving model Right grain Shared rules Additive facts 03 Query layer Filter, join Aggregate Push down 04 Presentation tool Navigation Display logic 05 Consumer BI · API · notebook AI assistant Ask, use, present Push work as far left as it can be done correctly and reusably Each layer assumes the one to its left has already resolved grain, duplicates and shared rules.
Shared computation lands in the serving model and query layer (highlighted). Consumers ask; they do not repair.

A customer-month serving model

Illustrative example

Instead of every report deriving customer-level figures from order lines, the sales data product publishes a customer-month serving model once. Revenue is declared additive; distinct customers is declared non-additive, so it is never summed across months.

The BI tool, the finance notebook and the AI assistant all query that model. Asked "what was average order value by region last quarter, in euros?", the assistant identifies the metric, dimension, period and currency. The query layer knows that average order value is a ratio, how quarters are defined and how currency is converted, and returns one answer. The assistant explains it.

In detail

Grain and additivity are modelling decisions

Model at the grain questions are asked, and classify every measure as additive, semi-additive or non-additive so that no engine or report has to rediscover it. A ratio is composed from sums, never summed.

A copy needs a reason

A new copy, materialisation or serving model should have a named justification: a consumer, a workload, a control or a resilience need. Otherwise, query the governed product where it already lives. This is governance in delivery applied to creating physical data.

Agnostic about tools

The direction does not choose a BI tool strategy. It constrains where shared computation lives, whatever tools sit on the right-hand side.

Not this

  • Not "materialise every KPI up front". Changeable definitions are composed at query time from stable facts.
  • Not one wide table for everything. Serving models keep the grain of the questions they serve.
  • Not a ban on presentation logic. Navigation and display belong in the tool.

Unresolved

  • Who authorises the justification for a new copy, and where is it recorded?
  • How are existing unjustified copies retired?
  • How should the query layer reject an invalid request, such as summing a non-additive measure, rather than silently returning a wrong answer?

Join the argument

Agree, challenge or add evidence

Architecture improves when its assumptions are challenged. Say which kind of contribution you're making:

  • Challenge I disagree because…
  • Evidence We've seen this too…
  • Question How would this work when…?
  • Alternative Another way to approach this…
  • Extension This also implies…

Please keep clients, employers and colleagues unidentifiable. Strong contributions may be quoted and credited on this site. Read the contribution guidelines and the privacy notice.

Comments are provided by giscus and stored on GitHub. Posting requires a GitHub account.

Next in the argument Process ledger fact Order, payment, shipment and return events are interpreted separately in every report, so net figures never quite reconcile. Semantic authority is independent of consumption