wunder beta

📘 How does bias hide inside data?

Statistical bias is the tendency of an estimation procedure to overestimate or underestimate the true population value in a consistent direction. Random error, by contrast, is unpredictable scatter that averages out over repeated measuremen

6
lessons
~30 min
to learn
Adults
level
Start the course →

What you’ll learn

  1. What Bias in Data Really MeansDefine statistical bias precisely and distinguish it from random error, variance, and everyday connotations of the word "bias."In statistics, bias is a systematic deviation that pushes estimates consistently away from the truth, and it is conceptually distinct from random error, which scatters estimates unpredictably around the truth. Crucially, large samples reduce random error but do not fix bias, so "more data" is not a cure. Throughout this course we treat bias as a property of a process, the way data is generated, collected, processed, and interpreted, not merely a property of a person's attitude. Naming where in that pipeline bias enters is the first skill of data literacy.
  2. Selection and Sampling Bias: The Literary Digest CaseDiagnose how non-representative sampling and self-selection distort conclusions, using the 1936 Literary Digest poll as a worked case.The 1936 Literary Digest poll is the canonical illustration of selection and nonresponse bias operating together to produce a catastrophic error. By drawing names from telephone, club, and subscriber lists, the magazine sampled wealthier voters, and by relying on voluntary mail-back responses it over-counted motivated, anti-incumbent voters. The result, predicting Landon over Roosevelt, was reversed by reality, and the episode reshaped survey methodology toward probability sampling. The enduring lesson is that representativeness, not raw size, determines whether a sample supports inference.
  3. Survivorship and Measurement BiasIdentify how missing-by-construction data (survivorship) and flawed instruments or definitions (measurement) silently corrupt conclusions.Survivorship bias arises when analysis is restricted to the units that passed through some selection filter, so the failures that would change the conclusion are invisible. Measurement bias arises earlier, when the instrument, proxy, or operational definition systematically mismeasures the construct of interest. Both are insidious because the corrupted data can look complete and clean. The defense is to ask explicitly what is missing and whether the numbers actually measure what they claim to.
  4. Aggregation Traps: Simpson's Paradox and the Berkeley CaseExplain how aggregating across a confounding variable can reverse an apparent trend, and use disaggregation to recover the truth.Simpson's paradox occurs when an association visible in aggregated data reverses or disappears once the data is split by a confounding subgroup. The 1973 UC Berkeley graduate admissions data is the classic case: in aggregate, men appeared admitted at a higher rate than women, yet department-by-department the pattern was parity or a slight edge for women, because women applied disproportionately to more competitive departments. The paradox is not a mathematical trick but a warning that the right level of analysis depends on the causal structure. Disaggregating along the confounder is the standard remedy.
  5. Algorithmic Bias: From Data to DecisionsMap where bias enters the machine-learning life cycle and connect data biases to real-world algorithmic harms.Machine-learning systems inherit and can amplify the biases in their data, so algorithmic bias is best understood as bias propagating through a life cycle from data generation to deployment. The Suresh and Guttag framework names distinct sources, historical, representation, measurement, learning, aggregation, evaluation, and deployment bias, that map cleanly onto the pipeline view introduced earlier. The Gender Shades audit and the ProPublica COMPAS investigation are documented cases showing how representation and label choices translate into disparate real-world outcomes. Mitigation requires intervening at the specific stage where the bias originates, not just tweaking the final model.
  6. Applied Project: Building Your Bias Audit ArtifactApply the course's concepts by producing a structured bias audit of a real dataset or published claim, naming each bias, its mechanism, and a concrete mitigation.In this capstone you will build a Bias Audit Brief: a short, evidence-based artifact that interrogates one real dataset or published quantitative claim using the pipeline framework. You will name each candidate bias by type, specify its mechanism and likely direction, and propose a stage-appropriate mitigation, all grounded in verifiable facts. The deliverable mirrors professional data-quality and model-card practice and becomes a portfolio piece for the Data Science & AI Foundations certificate. Rigor here means precise terminology, honest uncertainty, and refusal to assert any figure you cannot verify.

Questions this course answers

A digital thermometer consistently reads 0.5 degrees above the true temperature on every reading. This is best described as:

A consistent, same-direction offset is the signature of bias. Reproducibility (high reliability) does not make it accurate; the instrument is precisely wrong, which is exactly what bias means.

An analyst worries her survey estimate is biased, so she quadruples the sample size. What is the most likely effect?

Sample size governs random error, not bias. A larger sample yields a more precise estimate of the same systematically wrong value, which is why scaling data without fixing the process can deepen false confidence.

Which question is NOT useful for turning a vague claim of 'biased data' into an investigable hypothesis?

Dataset size addresses random error, not bias, so it does not help characterize a systematic deviation. The other three questions (reference, direction, mechanism) operationalize bias as a testable claim.

Why did the 1936 Literary Digest poll fail despite collecting about 2.4 million responses?

The frame (phone/auto/subscriber lists) skewed wealthy, and voluntary mail-back returns over-represented motivated opponents of Roosevelt. Both are systematic biases that size cannot fix; the count itself was not the problem.

Nonresponse bias is a concern specifically when:

A low response rate alone does not guarantee bias; the damage occurs when responders differ systematically from non-responders on the variable of interest, so response propensity is correlated with the outcome.

An online poll on a news site lets any reader vote on a political question. The biggest threat to its validity is:

Open online polls let the most motivated and unrepresentative voices include themselves, the modern form of the Digest's voluntary-response problem. Inclusion probability is uncontrolled, so the sample is not representative regardless of how many vote.

Grounded in trusted sources

  • OpenIntro Statistics, 4th Edition, Diez, Cetinkaya-Rundel & Barr (openintro.org/book/os)
  • U.S. Census Bureau, "Sources of Error" (census.gov), survey quality documentation
  • Squire, P. (1988). "Why the 1936 Literary Digest Poll Failed," Public Opinion Quarterly, 52(1), 125-133
  • OpenIntro Statistics, 4th Edition, chapter on data collection and sampling (openintro.org/book/os)
  • Wald, A. (1943/reprinted 1980). "A Method of Estimating Plane Vulnerability Based on Damage of Survivors," Center for Naval Analyses (work on WWII aircraft survivorship)
  • Groves, R. et al. (2009). Survey Methodology, 2nd Edition, Wiley (measurement error chapters)

Every Wunder lesson is built from real, reputable sources — never invented.

Related courses

Wunder is a personalized learn-anything platform — tell it any topic and it builds a beautiful, fact-checked course in minutes, with narration, a knowledge check, and a college-style University track.

All topics · Home

© 2026 Wunder Learning LLC · Terms & Privacy