📘 How do experiments prove a business change worked?
Causation, randomization, and guardrails—how experiments separate a real lift from a story.
What you’ll learn
- Why Experiment? Causation, Correlation, and the CounterfactualExplain why controlled experimentation is the most reliable way to establish that a business change causes an outcome rather than merely correlating with it.Business decisions hinge on causal claims, but observational data alone cannot separate the effect of a change from the confounders that travel with it. A controlled experiment manufactures a credible counterfactual by randomly assigning units to treatment and control, so the only systematic difference between groups is the change being tested. This is why randomized experiments are treated as the gold standard for causal inference, a framing made rigorous by R. A. Fisher's principles of randomization, replication, and blocking.
- Designing a Trustworthy Experiment: Hypotheses, Metrics, and PowerTranslate a business question into a testable hypothesis with a defined success metric, guardrails, and an adequately powered sample size before any data is collected.A trustworthy experiment is designed before it runs. You state a falsifiable hypothesis, choose an Overall Evaluation Criterion that captures success while guardrail metrics protect against hidden harm, and then use the baseline rate, a minimum detectable effect, the significance level, and the desired power to compute the required sample size and run time. Locking these decisions in advance is what prevents the analysis stage from drifting into wishful interpretation.
- Running Experiments and Reading Results HonestlyInterpret experiment results correctly by distinguishing signal from noise and recognizing the pitfalls that produce false or misleading conclusions.Even a well-designed experiment can be misread. Confidence intervals and p-values quantify uncertainty but are routinely misinterpreted, and practices like peeking, testing many metrics, and aggregating across imbalanced segments inflate false positives or reverse true effects. Sanity checks such as sample-ratio mismatch and a clear-eyed view of statistical versus practical significance separate trustworthy readings from convenient ones.
- When You Cannot Randomize: Quasi-Experiments and Their LimitsRecognize situations where randomized experiments are infeasible and apply quasi-experimental designs while understanding the stronger assumptions they require.Randomization is sometimes impossible for practical, ethical, or structural reasons, such as evaluating a price change, a regulation, or a network feature. Quasi-experimental designs, including natural experiments and difference-in-differences, recover causal estimates by exploiting comparison groups and timing rather than random assignment. They are powerful but rest on assumptions like parallel trends that, unlike randomization, cannot be guaranteed and must be argued and probed.
- Build It: A Mini Experiment Design from Question to DecisionProduce a complete, self-consistent experiment plan for a real business question, integrating hypothesis, metrics, power, analysis, and a decision rule into one artifact.This capstone walks you through building a one-page experiment design document, the artifact that experimentation teams use to make a test trustworthy before it launches. You will state the hypothesis, define the OEC and guardrails, set the unit of randomization, compute the required sample size from your MDE and power, pre-commit an analysis and decision rule, and choose between a randomized and a quasi-experimental approach when randomization is constrained. The result is a concrete, reusable template you can apply to any future business decision.
Questions this course answers
A manager observes that customers who use the new mobile app spend 30% more than those who do not, and concludes the app caused the increase. What is the central flaw?
App users chose to adopt the app, so they may already be more engaged or higher-spending. Without randomization, this confounding self-selection is indistinguishable from a true causal effect of the app.
Why is the fundamental problem of causal inference a problem at the level of the individual unit?
Each unit can only experience one condition at a time, so its counterfactual outcome is unobservable. Experiments address this by using a randomized control group as the estimated counterfactual for the treatment group.
What does randomization specifically accomplish that simply choosing a 'similar-looking' control group does not?
Hand-matching can only balance variables you know and measure. Random assignment balances both known and unknown confounders in expectation, which is its unique and central advantage.
A team sets significance at alpha = 0.05 and power at 0.80. What does the 0.80 represent?
Power is one minus the Type II error rate: the probability of correctly rejecting a false null, i.e., detecting a real effect of the assumed magnitude. Alpha, separately, governs the false-positive (Type I) rate.
Holding baseline rate, alpha, and power fixed, what happens to the required sample size if you halve the minimum detectable effect you want to catch?
Smaller effects are harder to distinguish from noise, so detecting them with the same power and significance requires substantially more units. This is why the MDE is set to a business-meaningful threshold rather than an arbitrarily tiny value.
Which is the best example of a guardrail metric rather than an OEC for a checkout redesign aimed at lifting purchases?
Page-load latency does not measure the experiment's success goal; it guards against a 'win' that degrades the user experience. The other options are plausible success or proxy metrics tied to the purchasing objective.
Grounded in trusted sources
- Ron Kohavi, Diane Tang, and Ya Xu, Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing (Cambridge University Press, 2020), ch. 1
- R. A. Fisher, The Design of Experiments (Oliver & Boyd, 1935)
- Ron Kohavi, Diane Tang, and Ya Xu, Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing (Cambridge University Press, 2020), chs. 2 and 17
- Jacob Cohen, Statistical Power Analysis for the Behavioral Sciences, 2nd ed. (Lawrence Erlbaum, 1988)
- Ron Kohavi, Diane Tang, and Ya Xu, Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing (Cambridge University Press, 2020), chs. 3 and 17-19
- American Statistical Association, 'ASA Statement on Statistical Significance and P-Values,' The American Statistician 70, no. 2 (2016): 129-133
Every Wunder lesson is built from real, reputable sources — never invented.
Related courses
Wunder is a personalized learn-anything platform — tell it any topic and it builds a beautiful, fact-checked course in minutes, with narration, a knowledge check, and a college-style University track.
© 2026 Wunder Learning LLC · Terms & Privacy