📊 Data Analysis
Turn raw data into insight. You'll clean and explore datasets, compute summaries, and build charts that reveal patterns and support honest conclusions.
What you’ll learn
- Data Analysis Changes the WorldShow, via Nightingale and Snow, that data analysis exists to change real decisions.Nightingale's 1858 rose diagram moved Parliament to reform army sanitation, and Snow's 1854 cholera map got a pump handle removed before germ theory existed. Both walked the same pipeline this course teaches: question, data, cleaning, exploration, conclusion, story.
- Start with a Sharp QuestionConvert vague prompts into sharp, measurable questions tied to decisions.A usable question names its metric, comparison, window, and the decision it feeds — pre-deciding what to measure, as Apollo mission control did, prevents 'data dredging,' where noise gets promoted to insight.
- Where Data Comes FromTrace data to its collection story and distinguish observational from experimental origins.Censuses, sensors, and logs each embed collection choices that bound what the data can say, and instruments drift just as respondents fib. The observational/experimental divide is fundamental: only randomized experiments straightforwardly support causal claims.
- Cleaning: The Unglamorous 80 PercentTreat cleaning — missing values, duplicates, standardization — as defensible analytical judgment.Real data arrives with defects, and preparation dominates analyst time; each fix (correct, standardize, flag, exclude) is a documented judgment call. The punched-card era already knew the rule: structure and validate at entry, because garbage in becomes garbage out at scale.
- Know Your Data TypesClassify variables (categorical, ordinal, discrete, continuous) and apply type-legal summaries.Data types are guardrails: means belong to numeric data, medians extend to ordinal, and categorical data allows only counts and modes — hence no 'average zip code.' Survey forms freeze types at design time, for better or worse.
- Center: Mean, Median, ModeChoose between mean, median, and mode based on skew and outliers.The mean chases extreme values while the median stays with the typical case — a $2M CEO makes the mean salary $272k in an office of $80k earners, and U.S. mean household income (~$114k) runs far above the median (~$80.6k). Skewed domains like incomes and markets demand the double check.
- Spread: Why Averages Lie AloneMeasure spread with range, IQR, and standard deviation, and explain why averages alone mislead.San Francisco and Wichita share an average temperature but not a climate — spread is half the story everywhere, from factory tolerances to delivery times. The IQR resists outliers and pairs with medians; standard deviation pairs with means on bell-shaped data.
- Distributions and HistogramsRead histograms and diagnose bell, skewed, and bimodal shapes.Binning values into a histogram exposes the distribution's whole shape at once: symmetric bells license means and standard deviations, long tails demand medians, and two humps mean two mixed populations that must be split before summarizing.
- Correlation and Scatter PlotsUse scatter plots and Pearson's r, respecting r's blindness to non-linear structure.Scatter plots reveal relationships — like the Preston curve linking national income to life expectancy — and r compresses linear association to one number from −1 to 1. Anscombe's quartet proves identical statistics can hide wildly different shapes: always plot first.
- Correlation Is Not CausationSeparate correlation from causation via confounders and Snow's interventional standard.Hidden common causes — summer heat behind ice cream and drownings, wealth behind chocolate and Nobels — make correlation the beginning of a question, not an answer. Snow's case became causal when anomalies fit the mechanism and removing the exposure stopped the outbreak.
- Sampling and BiasRecognize survivorship, selection, and nonresponse bias, and prefer random over large samples.Wald armored the bombers where survivors showed no holes; the Literary Digest's 2.4 million biased responses lost to Gallup's 50,000 representative ones in 1936. Sample size cannot cure bias — always ask who could not, or would not, end up in the data.
- Honest Charts and Dishonest OnesDetect and avoid chart distortions: truncation, cherry-picking, dual axes, area tricks.The Challenger engineers had the O-ring data but their charts never made the cold-temperature pattern visible — presentation is safety-critical. Truncated axes, cherry-picked windows, dual axes, and area distortions manufacture impressions; four skeptical checks catch most of them.
- Uncertainty: From Sample to ConclusionAttach margins of error to estimates and resist over-reading noise.Sample estimates wear a ± of roughly 1/√n — about 3 points for n = 1,000 — and precision depends on sample size, not population size. Official statistics are survey estimates; treating within-margin wiggles as trends is chasing noise.
- Telling the StoryCommunicate findings with the context-finding-evidence-so-what structure.The analyst's product is a changed mind: lead with the finding, prove it with one honest chart plus caveats, and name the decision. Nightingale's diagram and Du Bois's 1900 hand-drawn charts remain the masterclass in rigor joined to persuasive design.
- Capstone: Anatomy of a Great AnalysisIntegrate the full pipeline by re-running Snow's investigation end to end.Snow's study contains every modern stage: a sharp causal question, door-to-door collection, cleaning, the revealing map, anomaly checks that fit the mechanism, and an intervention that worked. That combination — not the map alone — is the standard every analysis should chase.
Questions this course answers
John Snow's 1854 cholera map was persuasive because it:
Plotting each death at its address revealed the tight cluster around the Broad Street pump — visual evidence strong enough to get the pump handle removed, decades before germ theory was accepted.
Florence Nightingale's rose diagram demonstrated that most Crimean War deaths came from:
Her month-by-month wedges showed disease deaths dwarfing combat deaths, driving sanitary reform of army hospitals.
Which is the sharpest version of a business question?
It names the metric (median session length), the comparison (before/after June 3), and the scope (existing users) — a stranger could compute the answer.
'Data dredging' refers to:
Searching without a prior question guarantees you'll find patterns — many of them pure noise. Sharp questions, set in advance, are the antidote.
The key advantage of experimental data over observational data is that experiments:
Random assignment breaks the link between treatment and hidden confounders, which is what licenses 'X causes Y' — observation alone shows association.
You find 'NY', 'N.Y.', and 'New York' in the same column. The right fix is to:
They're one category in three costumes; standardizing preserves the data while making counts and groupings correct.
Grounded in trusted sources
- John Snow, 'On the Mode of Communication of Cholera' (John Churchill, 1855)
- Edward Tufte, 'The Visual Display of Quantitative Information' (Graphics Press, 1983) and 'Visual Explanations' (1997)
- Darrell Huff, 'How to Lie with Statistics' (W. W. Norton, 1954)
- Charles Wheelan, 'Naked Statistics' (W. W. Norton, 2013)
- David Freedman, Robert Pisani & Roger Purves, 'Statistics', 4th ed. (W. W. Norton, 2007)
- Peverill Squire, 'Why the 1936 Literary Digest Poll Failed', Public Opinion Quarterly 52 (1988)
- U.S. Census Bureau — 'Income in the United States: 2023'
- Whitney Battle-Baptiste & Britt Rusert (eds.), 'W. E. B. Du Bois's Data Portraits: Visualizing Black America' (2018)
Every Wunder lesson is built from real, reputable sources — never invented.
Related Science courses
Wunder is a personalized learn-anything platform — tell it any topic and it builds a beautiful, fact-checked course in minutes, with narration, a knowledge check, and a college-style University track.
Browse more Science courses · All topics · Home
© 2026 Wunder Learning LLC · Terms & Privacy