The graph went up and to the right. The slide used the good green. Someone wrote “strong momentum” in the speaker notes.
Meanwhile, customers were learning how to avoid the thing we had optimised.
This is not a story about data being bad. It is a story about data being obedient. A metric will faithfully count what you ask it to count, including behaviour your product created for the wrong reason.
A number is a compressed story
Every metric hides a chain of judgment:
- Which behaviour counts?
- Whose behaviour counts?
- Over what period?
- What do we believe that behaviour represents?
- What harmful side effect would the metric fail to show?
By the time a number reaches a leadership dashboard, those choices have often disappeared. The metric looks objective because the argument has been compressed out of view.
Daily active users might represent recurring value. It might also represent repeated friction, compulsory usage, aggressive notifications or a workflow that now takes twice as many visits.
The number cannot tell you which story is true by itself.
When the proxy becomes the product
A useful metric is a proxy for value, not value itself. Trouble begins when the team starts changing the experience to move the proxy and stops checking whether the relationship still holds.
Imagine a collaboration product that measures messages sent. The team makes messaging easier, adds reminders and celebrates higher activity. But if customers wanted fewer coordination loops, more messages may mean the product is producing work rather than removing it.
This is why a North Star should express a belief about customer value and business value, not merely provide a large number. Amplitude’s account of changing its own North Star is useful precisely because it treats the metric as revisable when the product strategy and understanding of value change.
Changing the metric is not admitting analytics failed. Refusing to change it can be.
The organisational incentive hiding inside the chart
Metrics do not live outside politics. Teams are funded, praised and reorganised around them. Once a target becomes a career fact, caveats become socially expensive.
The analyst who says “activation increased, but only because we changed the definition” is not ruining the meeting. They are protecting the decision.
The PM who brings support volume beside conversion is not diluting focus. They are restoring the part of the system the headline number excluded.
Healthy metric reviews make room for three uncomfortable sentences:
- The number moved, but we do not yet know why.
- The number moved because our definition changed.
- The number moved while a balancing measure worsened.
Give every hero metric a suspicious friend
For any metric the team wants to increase, choose at least one balancing measure that could reveal harm.
If you optimise signup completion, watch early regret or cancellation. If you optimise time spent, watch task success and user sentiment. If you optimise feature adoption, watch whether the feature replaces an existing success path or merely adds clicks.
Then add qualitative evidence. A session replay, support conversation or interview does not overrule the aggregate. It helps explain what kind of behaviour the aggregate contains.
A metric system, not one sacred number
The North Star framework developed with John Cutler treats the headline metric as an expression of value supported by input metrics, not as a lonely scoreboard. That structure matters because teams need both a shared direction and diagnostic detail.
If the North Star moves, inputs can help explain where. If an input improves while the North Star does not, the assumed relationship may be weak. If both improve while customer complaints rise, a balancing measure is exposing value the model forgot.
This makes metrics a system of hypotheses. Each connection can be examined rather than inherited as truth.
Experimentation still needs a decision standard
The cross-company paper on practical online controlled experiments collects lessons from practitioners at organisations including Microsoft, Google, LinkedIn and Airbnb. Its existence is a reminder that trustworthy experimentation requires substantial craft: instrumentation, guardrail metrics, statistical judgment and an understanding of organisational context.
An experiment can estimate a causal effect on measured outcomes. It cannot decide which outcomes deserve priority or whether a short-term gain damages a longer relationship. Those remain product and ethical judgments.
Before celebrating a winning variant, ask whether the test population represents the people who will live with the decision, whether novelty affected behaviour, whether guardrails captured likely harm and whether the effect is meaningful enough to justify complexity.
Metric review as investigation
A useful review begins with predictions made before the result. What did the team expect to move, by how much and through which behaviour? Then compare the observed pattern. Surprises become learning rather than inconvenient noise.
Finish with a decision: continue, adapt, stop or investigate. A dashboard viewed without a decision question easily becomes office wallpaper.
Questions to take back to your team
- What customer behaviour does your headline metric claim to represent?
- When was that relationship last tested rather than repeated from an old strategy deck?
- Which balancing measure could reveal that the team is improving the number while damaging the experience?
- Who has enough permission to say “the metric rose, but this is not good news”?
- If the metric dropped tomorrow, which qualitative evidence would help you understand why?
Open the dashboard your team celebrates most. Can you explain the human behaviour inside every line, or has the graph become a story nobody is allowed to edit?
My take
The dangerous metric is not the imperfect one. All metrics are imperfect. The dangerous metric is the one an organisation is no longer allowed to question.
Do not ask only, “Did it go up?” Ask, “What behaviour produced this movement, for whom, and what else changed?”
The graph deserves celebration only after the story survives inspection.