📊 Statistics I
Learn to describe and reason about data honestly. You'll summarize distributions, understand probability basics, and interpret averages and spread without being fooled.
What you’ll learn
- You Are About to Be FooledRecognise that sample size governs how much a statistic varies, and adopt the course's central question: how easily could I have been fooled?Kahneman and Tversky's 1974 hospital problem shows that most people judge a small hospital and a large one equally likely to record days with over 60% boys, when small samples swing far more. The failure isn't a missing formula but a missing habit — a percentage from 15 births and one from 45 are different kinds of object even at the same value. Variation is not noise obscuring the answer; asking how much a number would bounce by chance alone is the engine of every tool in the course.
- The Average Is a Story With Most of It MissingDistinguish mean, median and mode, and use the gap between mean and median to detect skew and the choices it invites.An average is a compression, and the question is always what fell out. The mean uses every value, which lets a single billionaire drag the average income of a bar of ten $40,000 earners to over $90 million — a figure that describes nobody present — while the median stays at $40,000 because it counts only how many are above and below, never how far. When mean and median diverge, the distribution is skewed and somebody chose which number to report; income is always right-skewed, so 'average income rose' and 'the typical person got poorer' can both be true of the same data.
- Spread Is the Actual InformationExplain why spread is often the substantive answer, distinguish range, IQR and standard deviation, and apply Anscombe's lesson that summaries must be plotted.A river averaging one metre deep may be safe to wade or may drown you — the average is identical and everything that matters is in the spread. Range is honest but fragile, the interquartile range is robust like the median, and the standard deviation is best understood as the natural unit of surprise: one SD out is unremarkable, three is worth a second look. Anscombe's 1973 quartet shares mean, variance, correlation (0.816) and regression line (y = 3.00 + 0.500x) across four datasets that are a line, a curve, a line plus an outlier, and a relationship invented by a single point.
- Where Numbers Come FromExplain selection and non-response bias through the 1936 Literary Digest poll, and internalise that sample size cannot fix a biased sample.The Literary Digest mailed over 10 million cards, received 2.4 million back, and predicted Landon would win with 57.1%; Roosevelt took 60.8% and the electoral college 523–8, and the magazine folded in 1938. The mailing list was built from subscribers, automobile registrations and telephone directories — a precise picture of the comfortable fraction of Depression-era America, which leaned against Roosevelt — and non-response bias filtered it again, as those who most disliked Roosevelt were likelier to reply. Gallup polled a few thousand demographically representative people and called it correctly: size buys precision, bias determines what you are measuring at all.
- Randomness Does Not Look RandomRecognise that randomness produces clusters and streaks, and separate 'is the pattern real?' from 'does the pattern mean anything?' using Clarke's flying-bomb analysis.Random data is clumpy: in 20 fair coin flips a run of four or more appears about 77% of the time, and a run of five about half the time, which fuels both the gambler's fallacy and the hot-hand belief. Londoners in 1944 correctly observed that V-1 hits clustered and inferred targeting; R. D. Clarke divided 144 km² of south London into 576 quarter-km² squares, counted the 537 hits, and found the distribution matched the Poisson prediction almost exactly (chi-squared p = 0.88). The clusters were real and meaningless — exactly what randomness produces for free — and only computing what chance alone would give could reveal that.
- Why More Data Helps, and How SlowlyApply the law of large numbers, the standard error and the central limit theorem to explain what sample size buys and at what cost.Given an unbiased sample, more data buys precision, and the standard error — standard deviation divided by the square root of n — governs the price: halving uncertainty takes four times the data, and cutting it to a tenth takes a hundred times. Population size does not appear in the formula, which is why 1,000 people sample a nation about as well as a town — a cook tasting soup needs no bigger spoon for a bigger pot, provided it is stirred. The central limit theorem states that sample means are approximately normally distributed almost regardless of the underlying population, which is why the bell curve is a fact about averaging rather than about nature.
- What a Confidence Interval Actually SaysState correctly what a confidence interval means, and recognise that it accounts only for sampling error.A number without a range is not a measurement — a poll reporting 52% with a 95% interval of [49%, 55%] is compatible with a genuine minority, which the headline conceals. The 95% describes the procedure, not the interval: in frequentist statistics the true value is fixed rather than random, so it is the interval that varies from sample to sample, and about 95% of such intervals would contain the truth. The distinction earns its keep because the interval is blind to bias, bad questions and broken sampling frames — a Literary Digest poll can carry a tight, correct-looking margin of sampling error around a number 19 points wrong.
- The p-value, Stated CorrectlyState the p-value definition precisely and identify the four standard misinterpretations, including the confusion of significance with effect size.A p-value is the probability of data at least this extreme given that the null hypothesis is true — P(data | null) — not the probability that the null is true given the data, which is what people want and a different quantity entirely. It is therefore not the chance of a fluke, not a measure of effect size (a large enough sample makes trivial differences significant), and not the probability of replication. The p < 0.05 threshold is a convention Fisher offered in the 1920s as a rough line that researchers might vary by context, with no mathematical basis — yet journals and careers treat p = 0.049 and p = 0.051 as different events.
- The Garden of Forking PathsExplain multiple comparisons, the garden of forking paths and publication bias, and interpret the replication findings correctly.At p < 0.05, twenty independent tests of things that do nothing yield at least one false positive about 64% of the time — the xkcd jelly bean result is the method working as specified, and the error is reporting one test as though it were the only one run. Gelman's garden of forking paths describes the invisible version: dozens of individually defensible analysis choices made after seeing the data constitute a multiple-comparisons procedure with no visible count to correct, and publication bias then files the null results in drawers. Ioannidis argued in 2005 that most published findings are false, and the Open Science Collaboration's 2015 audit of 100 psychology studies found 97% of originals significant against 36% of replications, at half the effect size.
- Base Rates: The Most Expensive Thing You Don't KnowCompute how base rate governs the meaning of a positive result, and extend the principle to the interpretation of p-values.For a stipulated disease affecting 1 in 1,000 tested by a 99%-accurate test, screening 100,000 people yields 99 true positives against 999 false ones — so about 9% of positives are real, and both 'the test is 99% accurate' and 'nine of ten positives are wrong' are true simultaneously. The cause is that the 1% error rate applies to an enormous healthy population while true positives are drawn from a tiny sick one; raising prevalence to 1 in 10 moves the same test to 92% without changing its accuracy at all. The p-value has no slot for prior plausibility, which is why significance in a field testing long shots means something entirely different from significance in a field testing well-motivated hypotheses — Ioannidis's argument, restated.
- Correlation, Causation, and the Third ThingEnumerate the possible explanations for a correlation and explain why randomisation, uniquely, licenses causal claims.A correlation is compatible with X causing Y, Y causing X, a confounder Z causing both, or luck — and it contains no information about which, since correlation states only that two numbers move together. Observational associations routinely fail: supplement takers also exercise, eat well and see doctors ('the healthy user effect'), and hormone replacement therapy's observational cardiac benefit disappeared under randomisation. Randomisation works because a coin cannot correlate with any variable, measured or unmeasured — so it breaks the link between treatment and every confounder including those nobody has thought of, which statistical adjustment can never do.
- How Not to Be FooledConsolidate the course into a practical checklist and adopt calibration, rather than blanket scepticism, as the goal.Every chapter asked one question in a different costume: how easily could I have been fooled? The checklist follows — who got counted and who didn't, mean or median, what's the spread, how big is the effect rather than how significant, how many things were tried, what was the chance before I looked, and was anything randomised. The goal is calibration rather than blanket disbelief, which requires no work and preserves every prior belief: Gallup was right, Clarke's analysis calmed a city, and the 36% replication rate is known precisely because science audits itself in public.
Questions this course answers
Why does the smaller hospital record more days with over 60% boys?
A percentage from 15 births and a percentage from 45 births are not the same kind of object even when they're the same number. Getting 10+ heads in 15 flips is unremarkable; getting 27+ in 45 is genuinely rare. How much a number bounces cares enormously about sample size.
What does this course mean by saying 'the variation is the subject'?
The instinct is to treat movement in the data as mess sitting on top of the real answer. But the engine of every tool in the course is one question: how much would this bounce by chance alone? If the answer matches your exciting finding, you found nothing.
'Average income rose while the typical person got poorer.' Can both be true from the same data?
Neither statement is a lie; they answer different questions. Income has a floor at zero and no ceiling, so it's always right-skewed and the mean always exceeds the median. If you don't know which one you were handed, you don't know what you were told.
Why is the median unmoved when a billionaire enters a bar of ten people earning $40,000?
That indifference to distance is exactly what robustness means, and it's why income is almost always reported as a median. The mean uses every value, which is its strength and its weakness — one extreme value drags it anywhere.
Two towns both average 15 °C. Town A ranges 10–20 °C; Town B ranges −25 to +40 °C. What does this show?
The mean told you nothing actionable; the spread told you everything — one town needs a jacket, the other needs two wardrobes, buried pipes and freeze-thaw-rated roads. You don't care what a river averages. You care how deep it gets.
Anscombe's four datasets share the same mean, variance, correlation and regression line. What was his point?
All four pass an identical numerical check. Only your eyes reveal that one is a perfect curve, one is a line dragged by a lone outlier, and one has its entire 'relationship' manufactured by a single point. Anscombe argued for better practice, not less statistics.
Grounded in trusted sources
- David Spiegelhalter — The Art of Statistics: Learning from Data (2019)
- OpenIntro Statistics, 4th ed. (2019)
- F. J. Anscombe — The American Statistician 27(1) (1973)
- R. D. Clarke — Journal of the Institute of Actuaries 72 (1946)
- Peverill Squire — Public Opinion Quarterly 52(1) (1988)
- Open Science Collaboration — Science 349 (2015)
- John P. A. Ioannidis — PLOS Medicine 2(8) (2005)
- Wasserstein & Lazar — The American Statistician 70(2) (2016)
Every Wunder lesson is built from real, reputable sources — never invented.
Related Math courses
Wunder is a personalized learn-anything platform — tell it any topic and it builds a beautiful, fact-checked course in minutes, with narration, a knowledge check, and a college-style University track.
Browse more Math courses · All topics · Home
© 2026 Wunder Learning LLC · Terms & Privacy