wunder beta

📘 How do you run a usability test?

Usability testing means observing real, representative users as they attempt real tasks with your product, then watching where they struggle. The defining move is behavioral: you study what people DO, not what they SAY they would do. A user

7
lessons
~30 min
to learn
Adults
level
Start the course →

What you’ll learn

  1. Lesson 1: What Usability Testing Actually IsDefine usability testing as observing real users attempting real tasks and distinguish it from interviews and surveys.Usability testing is behavioral research: you watch representative users attempt concrete tasks and study what they do, not what they say. This exposes the gap between stated opinion and real behavior that surveys and interviews miss. Tests are formative (diagnostic, run during design to find and fix problems) or summative (evaluative, run to measure against benchmarks). A good test names WHY a task failed across categories like navigation, labeling, and recoverability. Its lane is interaction quality, not desirability or market demand.
  2. Lesson 2: Choosing a Method and FormatSelect an appropriate test format (moderated/unmoderated, in-person/remote, qualitative/quantitative) and recruit representative participants.Usability tests vary along several axes. Moderated tests allow live probing and suit early, ambiguous designs; unmoderated tests scale cheaply for well-defined tasks. In-person testing offers rich observation, while remote testing widens reach and ecological validity. Qualitative testing (a few users) explains why problems happen; quantitative testing (about 20 users, per NN/g) produces reliable metrics. Settings range from controlled labs to field studies to scrappy guerrilla tests. Above all, recruit participants who genuinely represent your users, using a screener and testing distinct segments separately.
  3. Lesson 3: How Many Users — The 5-User QuestionExplain the basis and limits of the '5 users find ~85% of problems' claim and decide when more users are warranted.The rule that ~5 users uncover ~85% of usability problems comes from Nielsen and Landauer's 1993 model, 1−(1−L)^n, where the per-user discovery probability L averaged about 0.31. Diminishing returns make three rounds of five (test-fix-retest) better than one round of fifteen. But 85% is an average of the currently discoverable problems, with real variance shown by critiques such as Faulkner (2003), and it assumes a homogeneous user group. Use more users for distinct segments (3-5 each), quantitative metrics (~20+), or high-risk domains.
  4. Lesson 4: Writing Realistic Tasks and ScenariosWrite realistic, non-leading task scenarios with pre-defined success criteria and validate them with a pilot.A usability task states a realistic goal and lets the user find the path; it never lists steps. Scenarios wrap tasks in a short, believable story that motivates natural behavior without revealing the answer. The biggest leading mistake is echoing the interface's own button labels, which tests reading instead of findability — use neutral, goal-oriented language. Define a concrete success criterion for each task in advance so outcomes are scored consistently. Sequence tasks to avoid spoilers, start with a warm-up, and always pilot the script before real sessions.
  5. Lesson 5: Running the Session — Think-Aloud & Avoiding BiasFacilitate a think-aloud session neutrally, avoiding the common moderator biases that contaminate findings.Think-aloud asks participants to verbalize their thoughts as they work, opening a window on their mental model; its roots are in Ericsson and Simon's 1980 verbal-report research, and Nielsen popularized it for discount usability testing. Concurrent think-aloud is rich but distorts timing, while retrospective preserves clean timing at the cost of memory. The moderator is the biggest threat to valid data: leading questions, premature help, and reactions to success all bias results. Neutral moves — bouncing questions back, the 5-second rule, reassuring users you test the product not them — keep the data honest.
  6. Lesson 6: Measuring & Prioritizing — Metrics, SUS, SeverityMeasure usability with task metrics and the SUS, then rate and prioritize issues by severity for action.Core metrics—task success rate, time on task, error rate, and assists—turn impressions into comparable evidence. The System Usability Scale (SUS), created by John Brooke in 1986, yields a single 0-100 score from ten alternating positive/negative items, scored by adjusting each item and multiplying the sum by 2.5; the average is about 68 and it is not a percentage. Nielsen's 0-4 severity scale, combining frequency, impact, and persistence, prioritizes issues. The deliverable is a skimmable, evidence-backed findings report that maps high-severity issues and quick wins to concrete fixes.
  7. Guided Project: Usability Test Plan + Findings Report (Applied Project)Produce a complete usability test plan and findings report for a real product flow, applying every prior lesson.This capstone walks you through producing two artifacts. The test plan defines behavioral research goals, justifies a format, specifies a screener and representative participants, and contains a piloted, non-leading task script with success criteria and an intro/consent preamble. You then run neutral think-aloud sessions, capturing timestamped data separated from interpretation. Finally you cluster observations into categorized issues, rate severity via frequency/impact/persistence, compute metrics (and optionally a correctly scored SUS), and assemble a skimmable findings report with an executive summary, evidence-backed issue cards, and prioritized recommendations.

Questions this course answers

A team surveys 500 users who rate their new dashboard 4.5 out of 5 for ease of use, yet support tickets about the dashboard keep climbing. What does this best illustrate about why usability testing is needed?

Usability testing's core value is observing what users DO, exposing the gap between stated opinion (a high survey rating) and actual behavior (rising tickets). Sample size is not the issue here, no rating threshold defines usability, and support tickets are a lagging signal, not a substitute for watching users attempt tasks.

A PM wants to decide, before launch, whether a redesigned signup flow has fewer friction points than the current one, fixing issues as they are found. Which describes this work?

Finding and fixing problems iteratively during design is the definition of formative (diagnostic) testing. Summative testing measures a finished product against benchmarks rather than feeding fixes back into design; surveys and interviews collect self-report rather than observed task performance.

Which question is OUTSIDE the proper scope of a usability test?

Market demand and profitability are questions of desirability and viability, answered by discovery research and business analysis, not by watching task interaction. The other three are all about whether and how people can use the design, which is exactly what usability testing measures.

You have an early, ambiguous prototype and want to understand WHY users misread its core concept, with freedom to ask follow-up questions live. Which configuration fits best?

Ambiguous, early designs benefit from moderated testing, where a facilitator can ask 'why?' and adapt to surprises. Unmoderated tests use a fixed script with no live probing; large quantitative studies measure rather than diagnose; surveys collect self-report, not observed interaction.

A team needs a statistically reliable task-success rate to report to leadership and track quarter over quarter. Roughly how many users does NN/g recommend for such quantitative measurement, and why?

Nielsen Norman Group recommends roughly 20 users for quantitative usability studies because reliable metrics require a larger sample. The famous '5 users' guideline applies to qualitative discovery of problems, not to producing trustworthy numbers; fixed figures like 100 or 1-2 are not the basis of the recommendation.

Why is recruiting your own teammates as test participants a poor choice for most usability tests?

Insiders carry prior knowledge of the design and the domain, so they navigate fluently past problems that confuse real, representative users — producing confident but misleading feedback. The issue is representativeness, not over-criticism, ethics, or consent.

Grounded in trusted sources

  • Nielsen Norman Group, 'Usability 101: Introduction to Usability' (nngroup.com)
  • Jakob Nielsen, Usability Engineering (Morgan Kaufmann, 1993)
  • Nielsen Norman Group, 'Quantitative vs. Qualitative Usability Testing' (nngroup.com)
  • Tom Tullis & Bill Albert, Measuring the User Experience, 2nd ed. (Morgan Kaufmann, 2013)
  • Nielsen, J. & Landauer, T. K., 'A Mathematical Model of the Finding of Usability Problems,' Proc. ACM INTERCHI '93 (1993)
  • Nielsen Norman Group, 'Why You Only Need to Test with 5 Users' (nngroup.com); Faulkner, L., 'Beyond the five-user assumption,' Behavior Research Methods (2003)

Every Wunder lesson is built from real, reputable sources — never invented.

Related courses

Wunder is a personalized learn-anything platform — tell it any topic and it builds a beautiful, fact-checked course in minutes, with narration, a knowledge check, and a college-style University track.

All topics · Home

© 2026 Wunder Learning LLC · Terms & Privacy