wunder beta

📘 How do you choose experiment metrics?

Every experiment ultimately reduces to a comparison of numbers, so the numbers you choose silently decide what you will and will

7
lessons
~30 min
to learn
Adults
level
Start the course →

What you’ll learn

  1. Why Metrics Decide What You BuildExplain why metric choice, as the operational definition of a goal, is often the highest-leverage decision in an experiment.A metric operationalizes a fuzzy goal into something measurable on every variant, so a wrong metric can lead a perfectly run experiment astray. Precise operational definitions surface hidden assumptions and resolve disagreements that are really about definitions. Because watched numbers get optimized, metric choice is partly incentive design. Decision metrics (few, trusted) must be separated from debugging metrics (many, exploratory), and every metric should be treated as a hypothesis to validate.
  2. The Metric HierarchyDescribe the goal/North-Star, driver, guardrail, and debugging tiers of a metric hierarchy and the role of each.Mature programs organize metrics into a hierarchy: a long-term North-Star at the top, driver/proxy metrics hypothesized to feed it, guardrails that must not regress, and debugging metrics that explain movements. The North-Star reflects realized user value but is usually too insensitive for a single short test, which is why more sensitive drivers exist. Guardrails such as latency and errors enforce 'do no harm.' Keeping the decision set small while allowing many diagnostics preserves trust.
  3. What Makes a Good MetricIdentify the properties—measurable, attributable, sensitive, hard to game, clearly directional, and consistently defined—of a trustworthy experiment metric.A good metric is measurable from real instrumentation, attributable to the treatment, and sensitive enough to move detectably given variance and sample size. It should be hard to game, meaning the cheapest path to move it also serves the goal, and have an unambiguous good direction. Definitions must be consistent across variants, with a clearly specified population, denominator, and time window. Sensitivity is a joint property of the metric and the experiment.
  4. Leading, Lagging, and Proxy MetricsDistinguish leading from lagging indicators and explain proxy metrics and the surrogate risk they carry.Lagging indicators measure the outcomes you ultimately care about but only after delay, while leading indicators move earlier and are believed to predict them, trading immediacy for fallibility. Proxies substitute for goals that are slow or impossible to measure directly, but the gap between proxy and goal creates surrogate risk: a treatment can move the proxy while harming the true goal. Proxies must be validated against the true goal across past experiments, and pairing quantity with quality narrows the gap.
  5. The Overall Evaluation Criterion (OEC)Define the OEC and explain how it must encode long-term value while remaining measurable in the short term.The OEC is the single agreed measure, possibly a weighted composite, used to decide whether a treatment is better overall. Per Kohavi, Tang and Xu, a trustworthy OEC must capture long-term value yet be measurable within a short experiment, leaning on validated drivers and leading metrics. Composite OECs force explicit trade-offs but risk arbitrary or gameable weights, so weights should be justified and stress-tested. The OEC must be pre-registered, and the real decision rule is 'OEC improves and no guardrail is violated.'
  6. Guardrails, Counter-Metrics, and Goodhart's LawState Goodhart's Law correctly and show how guardrails and counter-metrics defend against metric gaming.Goodhart's Law, in Strathern's 1997 phrasing, says that when a measure becomes a target it ceases to be a good measure, because optimization exploits the gap between metric and goal; the idea originates with economist Charles Goodhart. Guardrail metrics (latency, errors, complaints) are monitored as 'must not regress' constraints to catch harm. Counter-metrics pair with a primary metric to expose the harm of naively maximizing it. Since both teams and users exert optimization pressure, the design goal is a system hard to game without genuinely helping users.
  7. Case Study: Ratios, Simpson's Paradox, and ValidationApply correct ratio-metric statistics (the delta method), recognize Simpson's paradox in segmented results, and validate a metric system.Many metrics are ratios where the randomization unit differs from the denominator unit, so the naive variance formula understates uncertainty; the delta method gives a correct large-sample variance using both variances and their covariance. Simpson's paradox occurs when a trend within every segment reverses on pooling, often from differing traffic mix between variants, so segments must be inspected and populations balanced. Segmentation aids diagnosis but invites multiple-comparison false positives. Validation—A/A tests, past-experiment replay, degradation tests, and proxy-versus-truth checks—is continuous because relationships drift.

Questions this course answers

What is the best description of a metric in the context of experimentation?

A metric operationalizes a fuzzy goal into a concrete, computable quantity. It is not itself a test or a proof of causation, and choosing it well is distinct from running the experiment correctly.

Two teams report very different 'active user' counts for the same product. What is the most likely root cause?

Differences in what events count, the time window, and the population define 'active' differently. Many disagreements about metrics are really disagreements about definitions.

Why are decision metrics and debugging metrics kept separate?

A small, well-vetted set of decision metrics makes the ship/no-ship call, while a larger set of diagnostic metrics only explains movements. Mixing them leads to decisions based on noisy secondary numbers.

In a typical metric hierarchy, what role does a North-Star metric play?

The North-Star captures core long-term value. Because it is broad and long-term, it is often too insensitive to move in a single short experiment, which is why driver metrics exist.

Why are driver (proxy) metrics introduced beneath the North-Star?

Driver metrics are nearer-term, more sensitive measures hypothesized to causally feed the goal, so they can move within an experiment's window. Their causal link to the goal must be validated.

Which set best exemplifies guardrail metrics?

Guardrails are constraints you must not violate while improving something else; latency, errors, and complaints are classic examples that catch harm.

Grounded in trusted sources

  • Kohavi, Tang, and Xu, Trustworthy Online Controlled Experiments — OEC and guardrails
  • Goodhart’s Law discussions in measurement design (Strathern / popular treatments)
  • Google HEART / GSM metric frameworks
  • North Star / OMTM product analytics primers
  • Pearl / causal inference notes on Simpson’s paradox (applied to ratios)

Every Wunder lesson is built from real, reputable sources — never invented.

Related courses

Wunder is a personalized learn-anything platform — tell it any topic and it builds a beautiful, fact-checked course in minutes, with narration, a knowledge check, and a college-style University track.

All topics · Home

© 2026 Wunder Learning LLC · Terms & Privacy