Engineering
Metrics Are Build Artifacts, Not Raw Observations
Move beyond reading data to owning the logic.
May 29, 2026 · Engineering · Leon Liang
The common belief is that data teams earn their strategic seat by learning to read the business. The logic follows that if you understand same-store sales, cohort retention, and the CFO’s priorities, you graduate from a ticket-taker to a partner.
This is only half right. Reading a metric is table stakes. The real work is owning the definition. The number on a dashboard is not a raw observation of the world; it is an artifact produced by a pipeline, designed by someone’s specific logic.
Frameworks are looser than their retellings
Before debating which metrics matter, it is worth noting that the industry canon is rarely as rigorous as the diagrams suggest.
The metrics canon is mostly forty to fifty years old
-
1975
Goodhart's law
An aside in a monetary-policy paper, not a management principle.
-
1983
Andy Grove documents OKRs in High Output Management
A few pages describing how Intel already worked.
-
2007
Dave McClure presents AARRR
A conference slide deck, later hardened into a funnel.
-
2009
Eric Ries names vanity metrics
-
2010
Google publishes HEART at CHI
The one entry here that is a peer-reviewed paper.
-
2017
Amplitude systematises the North Star Metric
Sean Ellis coined it years earlier as a looser heuristic.
The HEART framework is the outlier here. Published as actual research at CHI 2010 by Rodden, Hutchinson, and Fu, its primary value is not the five categories, but the “Goals-Signals-Metrics” process. It forces teams to decide what they want to achieve and what observable behavior indicates success before deciding what to measure. Most teams skip the first two steps and argue over the third.
Other frameworks are lighter than their reputations. AARRR began as a 2007 deck where Dave McClure urged startups to “concentrate on stuff that really matters,” not to follow a rigid linear funnel. The North Star Metric was Sean Ellis’s heuristic for early-stage companies before Amplitude’s 2017 playbook systematized it for the enterprise. Even OKRs are often credited to John Doerr and Google, though Doerr maintains he was simply the messenger for Andy Grove’s practices at Intel, documented in High Output Management (1983).
These frameworks are not useless; they are heuristics. Recognizing this lowers the temperature in meetings where someone insists a framework is being applied “incorrectly.”
The misquoted Goodhart’s Law
The most cited version of Goodhart’s Law—“when a measure becomes a target, it ceases to be a good measure”—is actually a 1997 paraphrase by Marilyn Strathern. It is catchier and broader than the original, which was a specific observation about statistical regularities collapsing once policymakers acted upon them.
This distinction is practical. The popular version implies that measurement is inherently self-defeating, which leads to despair. The original implies something more useful: a relationship that held while ignored may not survive being optimized. This is an argument for monitoring whether your metric still correlates with your desired outcome, not an argument against setting targets.
The number is an artifact with a build process
Frameworks rarely address the implementation. Whatever metric you choose, a system must compute it. In most organizations, that “system” consists of several different tools disagreeing quietly.
A semantic layer is the modern solution. Unlike another dashboard, dbt’s Semantic Layer and MetricFlow move the metric definition out of the BI tool and into a modeled, version-controlled layer. Changing a definition updates it everywhere it is invoked, rather than relying on a single dashboard owner to remember the change.
The engineering shift is the real win. When revenue is defined once and join logic is generated deterministically, two dashboards cannot silently disagree. The metric becomes a reviewable object in a pull request, complete with a diff, a blame history, and a test. This is fundamentally different from a number typed into a chart configuration.
Focus on input metrics
Amazon’s internal practice, documented in Working Backwards by Colin Bryar and Bill Carr, offers a critical distinction: separate controllable input metrics from output metrics.
Revenue is an output. No one can “do” revenue on a Monday morning. However, the inputs that drive it—selection breadth, price competitiveness, page speed, and in-stock rates—are things a team can act on directly. The discipline is to instrument the inputs and let the output follow.
This connects back to the artifact argument. Output metrics are usually well-defined because everyone watches them. Input metrics are where definitions rot; they are often owned by a single team, rarely audited, and quietly redefined when an upstream model changes. They are the metrics you can actually move, yet they are the least likely to be trustworthy.
The “Single Source of Truth” myth
For those launching metrics-unification programs, Will Kelly’s September 2025 essay is essential reading. He argues that Finance and Engineering will always maintain competing versions of the truth because they answer different questions under different constraints.
The goal should be to surface disagreements rather than eliminate them. A documented, deliberate difference between a finance number and a product number—with both definitions and the reasoning visible—is healthier than a single “blessed” number that half the organization privately distrusts.
Aeolus view — If two dashboards disagree, the problem is rarely the dashboards. The problem is that the metric has no owner, no version-controlled definition, and no test. It isn’t actually a metric yet. Before adopting a new framework, identify your five most-cited numbers and find the single place their logic lives. If you can’t find it, that is your project. It is unglamorous and takes weeks rather than quarters, but it builds more business trust than any amount of dashboarding.
The trust ceiling
dbt’s 2025 State of Analytics Engineering report (April 2025, 459 respondents) found that 75% of respondents believe their organizations highly value and trust their data teams. While this is a selection-biased sample of teams already using modern tooling, it suggests the trust ceiling is higher than many data teams assume.
The teams that reach that ceiling are rarely the ones with the most sophisticated models. They are the ones whose numbers hold up under scrutiny.
The bottom line
Learning to read the business is valuable; a data engineer who can interrogate an earnings report is an asset. But the durable version of that skill is not interpretation. It is being the person who can explain exactly how a number is produced, show where the logic lives, and change it in one place.
If your metrics are scattered across seven dashboards and no one knows which is authoritative, we can help. Often, the solution is a fortnight of definition work rather than a full platform migration. We would rather tell you that.
Want a second opinion on your data stack?
Every Aeolus engagement starts with a fixed-fee data & AI-readiness audit — a short, low-risk first step before any larger build.
Book a data & AI-readiness audit