Skip to content

AI-ready data foundations · Production DataOps & pipeline engineering · Fractional data leadership · CI/CD for data pipelines · Schema drift, caught before it ships · RAG-ready semantic layers · Infrastructure as code · North America, remote-first

← All insights

Engineering

A metric is a build artifact, not a number

Reading the number is the easy part.

May 29, 2026 · Engineering · Aeolus Data

A copperplate engraving of an apothecary's beam balance in equilibrium, hatched in deep navy on cream paper, with a graduated row of precision weights on the bench beside it and the smallest weight tinted pink and set apart from the set.

There is a widely held belief that a data team earns its strategic seat by learning to read the business. Understand same-store sales, understand cohort retention, understand what the CFO actually looks at, and you graduate from ticket-taker to partner.

That belief is half right, and the missing half is the part that pays. Reading a metric is table stakes. Owning its definition is the job, because the number on the dashboard is not an observation about the world. It is an artifact your pipeline produced, and somebody decided how.

The frameworks are older and looser than their retellings

Before arguing about which metrics matter, it helps to know that most of the canon was never as rigorous as the diagrams suggest.

The metrics canon is mostly forty to fifty years old

  1. 1975

    Goodhart's law

    An aside in a monetary-policy paper, not a management principle.

  2. 1983

    Andy Grove documents OKRs in High Output Management

    A few pages describing how Intel already worked.

  3. 2007

    Dave McClure presents AARRR

    A conference slide deck, later hardened into a funnel.

  4. 2009

    Eric Ries names vanity metrics

  5. 2010

    Google publishes HEART at CHI

    The one entry here that is a peer-reviewed paper.

  6. 2017

    Amplitude systematises the North Star Metric

    Sean Ellis coined it years earlier as a looser heuristic.

Goodhart 1975; Grove, High Output Management, 1983; McClure, Startup Metrics for Pirates, 2007; Ries via Tim Ferriss's blog, 19 May 2009; Rodden, Hutchinson and Fu, CHI 2010; Amplitude, North Star Playbook, 2017.

The HEART framework is the outlier in that list, and the one worth reading in the original. Rodden, Hutchinson and Fu published it at CHI 2010 as actual research, and its most useful contribution is not the five categories but the Goals-Signals-Metrics process that precedes them: decide what you are trying to achieve, then what observable behaviour would indicate it, and only then what to measure. Most teams skip the first two steps and argue about the third.

The others are lighter than their reputations. AARRR began as a 2007 deck in which Dave McClure wanted startups “concentrating on stuff that really matters,” not a rigid linear funnel. The North Star Metric started as Sean Ellis’s one-metric-that-matters heuristic for early-stage companies and was systematised into an enterprise framework by Amplitude’s playbook in 2017. OKRs are routinely credited to John Doerr and Google, when Doerr has consistently said he was the messenger and Andy Grove the inventor, having documented the practice at Intel in High Output Management in 1983.

None of this makes the frameworks useless. It makes them heuristics, which is how they were offered, and it should lower the temperature of any meeting where someone insists a framework is being applied incorrectly.

Goodhart’s law is not what almost anyone quotes

The line everyone reaches for is “when a measure becomes a target, it ceases to be a good measure.” That phrasing is not Goodhart’s. It is Marilyn Strathern’s 1997 paraphrase, and it is both catchier and broader than the original, which was a narrower observation about statistical regularities collapsing once policymakers act on them.

The distinction matters in practice. The popular version implies measurement is inherently self-defeating, which counsels despair. The original implies something more useful: a relationship that held while nobody was optimising for it may not survive being optimised for. That is an argument for watching whether your metric still correlates with the outcome you care about, not an argument against targets.

The number is an artifact, and it has a build process

Here is the part the frameworks do not address at all. Whatever metric you choose, some system has to compute it, and in most organisations that system is several systems disagreeing quietly.

A semantic layer is the current answer, and it is a genuinely different proposition from another dashboard. dbt’s Semantic Layer and MetricFlow move the metric definition out of the BI tool and into the modelled, version-controlled layer, so that changing a definition changes it everywhere it is invoked rather than in the one dashboard whose owner remembered.

The engineering consequence is the interesting one. If revenue is defined once and the join logic is generated deterministically from that definition, then two dashboards cannot silently disagree, because there is only one implementation. The metric becomes reviewable in a pull request. It gets a diff, a blame history, and a test. That is a materially different object from a number someone typed into a chart configuration.

Input metrics are the ones you can actually move

The most useful distinction we know of comes from Amazon’s internal practice, documented in Working Backwards by Colin Bryar and Bill Carr: separate controllable input metrics from output metrics. Revenue is an output. Nobody can go to work on Monday and do revenue. The inputs that produce it, selection breadth, price competitiveness, page speed, in-stock rate, are things a team can act on directly, and the discipline is to instrument those and let the output follow.

This maps cleanly onto the artifact argument. Output metrics tend to be well defined and widely agreed, because everyone looks at them. Input metrics are where definitions rot, because each is owned by one team, checked by nobody, and quietly redefined when someone changes an upstream model. They are simultaneously the metrics you can move and the metrics least likely to be trustworthy.

A single source of truth is mostly a fairytale

Worth reading if you are about to launch a metrics-unification programme: Will Kelly’s essay on the single source of truth, published in September 2025, argues from experience that Finance and Engineering will maintain competing versions of the truth no matter how much unification you attempt, because they are answering different questions under different constraints.

His prescription is the one we would give too. Aim to surface disagreements rather than eliminate them. A documented, deliberate difference between the finance number and the product number, with both definitions visible and the reason recorded, is a healthier system than a single blessed number that half the organisation privately does not believe.

Aeolus view. If two dashboards disagree, the problem is almost never the dashboards. It is that the metric has no owner, no definition in version control, and no test, so it is not really one metric. Before adopting a framework, pick your five most-cited numbers and try to find, for each, the single place its logic lives. If you cannot, that is the project. It is unglamorous, it takes a few weeks rather than a few quarters, and it does more for how the business trusts you than any amount of dashboard work.

What this is worth

dbt’s 2025 State of Analytics Engineering report, published in April 2025 from 459 respondents, found 75 percent agreeing that their organisations highly value and trust their data teams. That is dbt’s own survey of dbt’s own community, so it is a selection-biased sample of teams already invested in modern tooling, and we would not lean on the number as an industry base rate.

What it does suggest is that the trust ceiling is not as low as data teams often assume. The teams that get there are rarely the ones with the most sophisticated models. They are the ones whose numbers hold up when someone checks.

Where this leaves you

Learning to read the business is worth doing, and any data engineer who can interrogate an earnings report is more useful than one who cannot. But the durable version of that skill is not interpretation. It is being the person who can say exactly how a number is produced, show where that logic lives, and change it in one place.

If your metrics live in seven dashboards and nobody is sure which is authoritative, we are happy to look at it with you. Often the answer is a fortnight of definition work rather than a platform migration, and we would rather tell you that.

Want a second opinion on your data stack?

Every Aeolus engagement starts with a fixed-fee data & AI-readiness audit — a short, low-risk first step before any larger build.

Book a data & AI-readiness audit