📘 How do you power an experiment?
Statistical power is the probability that a test correctly rejects the null hypothesis when a true effect exists. Formally, power equals
What you’ll learn
- Why Power Matters in ExperimentsExplain what statistical power is, how it relates to Type I and Type II errors, and why it must be planned before data collection.Statistical power is the probability of detecting a true effect, equal to 1 minus beta (the Type II error rate). Hypothesis tests risk two errors: false positives at rate alpha and false negatives at rate beta. Underpowered studies waste resources by missing real effects and inflating those they do detect. Power is a design choice fixed before launch, not a quantity to compute after the fact. In online experimentation, power determines which often-tiny business effects can be reliably measured.
- The Four Interlocking QuantitiesShow how alpha, power, effect size, and sample size are mathematically linked so that fixing any three determines the fourth.Sample size planning revolves around four interdependent quantities: significance level alpha, power (1 - beta), effect size, and sample size. Fixing three determines the fourth, typically solving for sample size given alpha, power, and a target effect. Power increases with larger samples, larger effects, and looser alpha, and decreases with greater variance. The non-centrality parameter of the test distribution formalizes this, growing with the square root of sample size. The 0.05/0.80 conventions are defaults that should be matched to the costs of each error type.
- Effect Size and Baseline VarianceDefine standardized effect size (Cohen's d), explain Cohen's benchmarks, and show how baseline variance affects the signal-to-noise ratio.Effect size expresses a difference in scale-free terms; Cohen's d divides the mean difference by the pooled standard deviation. Cohen's rough benchmarks are 0.2 (small), 0.5 (medium), and 0.8 (large), offered as last-resort defaults. Because variability sits in the denominator, noisier outcomes require larger samples for the same mean difference. For binary outcomes, the baseline rate p sets the variance p(1 - p), and effects are often stated as relative lift. Inputs for planning come from pilot data and historical baselines, both carrying uncertainty.
- Minimum Detectable Effect and Sample SizeDefine the minimum detectable effect and derive how required sample size scales inversely with the square of the MDE.The minimum detectable effect is the smallest true effect an experiment can reliably detect given its alpha, power, and sample size. Required sample size is approximately proportional to variance divided by the square of the MDE, so halving the MDE roughly quadruples the needed sample. The standard two-sample formula uses the constant (z for alpha/2 plus z for beta) squared, about 7.85 at alpha = 0.05 and 80% power. Inverting the formula, the MDE falls only with the square root of sample size, setting realistic expectations from fixed traffic.
- Pitfalls: Practical Significance and the Winner's CurseDistinguish statistical from practical significance and explain how low power causes missed effects, magnitude inflation, sign errors, and the winner's curse.Statistical significance means an effect is unlikely under the null, not that it is large enough to matter; with big samples trivial effects become significant, so judge results against a pre-set practical threshold using confidence intervals. Low power directly raises Type II errors, so real effects are missed and findings fail to replicate. When an underpowered study does reach significance, the estimate is systematically inflated (a Type M error) because only the largest-by-chance results clear the bar. In online experiments this is the winner's curse: winning treatments overstate their lift and disappoint after launch. At very low power, results can even carry the wrong sign (a Type S error), and adequate power remedies all of these at once.
- Duration Planning and Variance ReductionTranslate required sample size into test duration and show how variance reduction techniques like CUPED shrink the needed sample.Dividing the required sample by daily eligible traffic gives the minimum duration, adjusted for exposure fraction, allocation, and concurrent experiments. Experiments should span full business cycles, at least one or two weeks, and use a pre-committed stopping rule to avoid peeking bias. CUPED, introduced by Deng and colleagues in 2013, uses pre-experiment data as a covariate to lower variance without bias, with reduction roughly the square of the covariate-outcome correlation. Other levers, such as less noisy metrics, stratification, winsorizing, and paired designs, further shrink the required sample.
Questions this course answers
Statistical power is defined as which of the following?
Power is the probability of correctly rejecting a false null, equal to 1 minus the Type II error rate beta. Alpha governs false positives, not power.
If an experiment is designed with 80% power, what is the probability it misses a true effect of the assumed size?
Beta = 1 - power = 1 - 0.80 = 0.20, so there is a 20% chance of a Type II error (missing the effect) at the assumed effect size.
Why is post-hoc power computed from the observed effect generally discouraged?
Observed (post-hoc) power is a deterministic function of the p-value, so it provides no new information and is often misinterpreted. Power should be planned before data collection.
In a standard power analysis, how many of the four quantities (alpha, power, effect size, sample size) must be fixed to determine the remaining one?
The four quantities are mathematically linked; fixing any three determines the fourth. Typically alpha, power, and effect size are set to solve for sample size.
Holding everything else constant, which change DECREASES statistical power?
Larger outcome variance makes the effect harder to detect, lowering power. Bigger samples, bigger effects, and a looser alpha all raise power.
The conventions alpha = 0.05 and power = 0.80 are best described as:
These are widely used conventions, not laws. The appropriate values depend on the relative costs of false positives versus false negatives in the specific decision context.
Grounded in trusted sources
- Kohavi, Tang, and Xu, Trustworthy Online Controlled Experiments — power and MDE
- Cohen, Statistical Power Analysis for the Behavioral Sciences
- Gelman & Carlin, on Type M / Type S errors and the winner’s curse
- Statsig / Optimizely / industry primers on experiment duration and variance reduction
- OpenIntro / ISL chapters on hypothesis tests and power
Every Wunder lesson is built from real, reputable sources — never invented.
Related courses
Wunder is a personalized learn-anything platform — tell it any topic and it builds a beautiful, fact-checked course in minutes, with narration, a knowledge check, and a college-style University track.
© 2026 Wunder Learning LLC · Terms & Privacy