wunder beta

📘 How does hypothesis testing decide evidence?

We almost never observe an entire population; we observe a sample and must reason backward to the population that produced it. Hypothesis testing formalizes this by stating a specific, testable claim about a population parameter (such as a

6
lessons
~30 min
to learn
Adults
level
Start the course →

What you’ll learn

  1. The Logic of Hypothesis TestingExplain why hypothesis testing frames inference as a falsifiable claim about a population and distinguish the null from the alternative hypothesis.Hypothesis testing is a structured method for deciding whether sample data are surprising enough to cast doubt on a default claim about a population. We assume the null hypothesis is true, then ask how compatible the observed data are with that assumption. The framework cannot prove a hypothesis true; it can only fail to reject, or reject, the null. This asymmetry, rooted in Fisher's notion that a null is 'never proved but possibly disproved,' is the conceptual core of everything that follows.
  2. Errors, Significance Level, and PowerDefine Type I and Type II errors and relate the significance level, power, effect size, and sample size to the design of a test.Because we decide from limited data, every test risks two mistakes: rejecting a true null (Type I) and failing to reject a false null (Type II). The significance level alpha caps the Type I rate we are willing to tolerate, while beta is the Type II rate and power (1 minus beta) is the chance of detecting a real effect. These quantities trade off against one another and depend on effect size and sample size, so a sound test is designed before data are collected.
  3. Test Statistics, Sampling Distributions, and the p-valueCompute and interpret a test statistic and p-value relative to a sampling distribution, and correctly state what a p-value does and does not mean.A test statistic standardizes how far the observed data fall from what the null predicts, and its sampling distribution under the null tells us how often such extremes arise by chance alone. The p-value is the probability, computed assuming the null is true, of getting a result at least as extreme as the one observed. We reject the null when the p-value is below alpha, but the p-value is not the probability the null is true and, by itself, is a limited measure of evidence.
  4. Choosing and Running the Right TestSelect an appropriate hypothesis test for a given research question and data type, check its assumptions, and carry out the procedure.Different questions and data types call for different tests: one-sample and two-sample t-tests for means, z-tests and chi-square tests for proportions and categorical data, among others. Each test rests on assumptions — about independence, the scale of measurement, and the distribution of the data — that must be checked, because violating them distorts the error rates. Choosing the test and the tail (one- or two-sided) is part of the pre-registered analysis plan, not a post-hoc decision.
  5. Interpreting Results, Pitfalls, and ReproducibilityInterpret test results alongside confidence intervals and effect sizes while recognizing and avoiding common abuses such as p-hacking and uncorrected multiple comparisons.A responsible interpretation reports more than a single p-value: it pairs the decision with a confidence interval, an effect size, and full transparency about the analyses run. Many published failures trace to predictable pitfalls — testing many hypotheses without correction, stopping data collection when significance appears, and treating 0.05 as a bright line. Understanding these abuses, and the safeguards against them, is what separates valid inference from spurious findings.
  6. Guided Project: Build a Hypothesis-Testing Mini ArtifactApply the full workflow end to end by designing, executing, and documenting a complete two-group hypothesis test on a real dataset.In this capstone you build a reproducible 'mini artifact' — a documented analysis that takes one research question from hypotheses through assumption checks, test execution, and an honest interpretation. You will work with a real, openly licensed dataset, pre-specify your plan, run an appropriate two-group test, and report the p-value together with an effect size and confidence interval. The deliverable demonstrates every principle from the prior lessons in a single auditable workflow.

Questions this course answers

Which statement correctly characterizes the null hypothesis?

The null is a statement about a population parameter (not the sample statistic) that we provisionally assume true. Fisher noted it can be disproved but never proved, so consistency with the data does not establish it as true.

Why must the direction of a one-sided alternative hypothesis be chosen before examining the data?

Selecting the direction after seeing which way the data point doubles the effective chance of a false positive and is a form of bias; the alternative direction (and the test as a whole) should be specified in advance.

After a test, a researcher reports 'we fail to reject H0.' What is the most accurate interpretation?

Failing to reject means the evidence was insufficient to overturn the default; it never confirms the null. Absence of evidence against H0 is not evidence that H0 is true.

A Type I error occurs when:

A Type I error is a false positive: rejecting H0 when it is in fact true. Failing to reject a false null is a Type II error.

For a fixed sample size, what is the effect of lowering the significance level alpha from 0.05 to 0.01?

A smaller alpha pushes the critical value further into the tail, shrinking the false-positive rate but enlarging beta, which lowers power. With n fixed, the two error rates trade off.

Power is best defined as:

Power is 1 minus the Type II error rate: the chance of detecting a real effect. It increases with larger effect size, larger sample size, and higher alpha.

Grounded in trusted sources

  • Fisher, R. A. (1935). The Design of Experiments. Oliver and Boyd.
  • Wasserman, L. (2004). All of Statistics: A Concise Course in Statistical Inference. Springer. Chapter 10, Hypothesis Testing and p-values.
  • Neyman, J., & Pearson, E. S. (1933). On the Problem of the Most Efficient Tests of Statistical Hypotheses. Philosophical Transactions of the Royal Society of London, Series A, 231, 289–337.
  • Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences (2nd ed.). Lawrence Erlbaum Associates.
  • Wasserstein, R. L., & Lazar, N. A. (2016). The ASA's Statement on p-Values: Context, Process, and Purpose. The American Statistician, 70(2), 129–133.
  • Moore, D. S., McCabe, G. P., & Craig, B. A. (2017). Introduction to the Practice of Statistics (9th ed.). W. H. Freeman.

Every Wunder lesson is built from real, reputable sources — never invented.

Related courses

Wunder is a personalized learn-anything platform — tell it any topic and it builds a beautiful, fact-checked course in minutes, with narration, a knowledge check, and a college-style University track.

All topics · Home

© 2026 Wunder Learning LLC · Terms & Privacy