Test your app
Website Usability Testing: A 3-Task Script You Can Run This Week
Website usability testing without a lab: write 3 tasks, ask testers to think aloud, record what they do, and rate what you find by severity.
Website usability testing means watching real people try to do real tasks on your site while you stay quiet and take notes. Write 3 tasks, ask each tester to think aloud, record what they do (not just what they say), and rate each problem by how badly it blocks the task.
You built the site, so you can no longer see it the way a first-time visitor does. A usability test gives you that view back. It needs a short script, a few kind people, and the willingness to sit on your hands while they work.
What is website usability testing, and what is it not?
Nielsen Norman Group (NN/g) describes a usability-testing session as a researcher asking a participant to perform tasks, usually on one or more specific interfaces (NN/g, page last reviewed 15 July 2026, read 6 October 2026). Think of it as a trio: a facilitator who guides the session, tasks that resemble real life, and a participant who resembles a real user. That page also says small wording slips in a task can change how people behave, so we write ours with care.
So it tests the site, not the person. It is also not a poll. "Do you like it?" collects opinions. A task collects behavior, and behavior is what you can fix.
For who to invite, what to ask, and how to run the 5-person round, our guide on getting honest website feedback has it covered, including how to think aloud with a tester. This guide is about the script itself, what to write down, and how to rank what you find.
How do you write tasks instead of opinions?
A good task is something a visitor would actually want to do, written as a short situation. NN/g's guidance on task scenarios says they should be realistic, push the person to take an action, and avoid leading, meaning they should not give away how the interface is meant to be used (NN/g, 12 January 2014, read 6 October 2026).
Here is the difference:
- Opinion prompt: "What do you think of the menu?"
- Leading task: "Click Events, then find the Saturday story hour."
- Good task: "You want something to do with a 4-year-old this Saturday morning. Find one thing you could go to."
The good task never names a button. If the tester has to hunt for the path, you learn whether the path is findable. That is the whole point.
What does a 3-task script look like?
Use 3 tasks, in this order:
- Find. The visitor looks for something. This tests whether the site makes its main content easy to locate.
- Do. The visitor completes the one action your site exists for, such as signing up, booking, or buying.
- Undo or recover. The visitor cancels, changes, or fixes a mistake. We like it because a visitor who can't back out of something stops trusting the page.
We suggest printing the script and putting each task on a card the tester reads aloud, so the wording stays the same for everyone. Ask testers to use their own device, and make sure at least one is on a phone, because a layout that works on your laptop can fail on a small screen. That is also why our audit reads the rendered page plus desktop and mobile screenshots.
If you are curious how this lines up with our six areas: Do and Undo cover much of the same ground as Flow and Usability, which looks at what happens when someone signs up, buys, or hits a dead end. The phone is where Mobile and Access comes in.
What should you record?
Record what happened, not your verdict about it. For each task, log:
- Result: finished alone, finished with a nudge, gave up, or did the wrong thing and believed it was right.
- Time: roughly how long, using the clock on your phone.
- Path: the pages and clicks they used, including detours.
- Words: one or 2 exact phrases that show what they expected, written as said.
- Surprises: anything you did not predict.
If the tester agrees, record the screen and audio so you can check your notes later. Recording is optional help, and a phone's screen recorder does the job at our scale. Ask permission first, and say what you will do with the recording.
The "did the wrong thing and believed it was right" result deserves attention. A visitor who reports success after failing is more worrying than one who gives up, because nothing on the page told them.
How do you rate what you find?
Rating keeps you from fixing the loudest problem first. NN/g's severity page, written in 1994 for evaluators judging problems in a heuristic evaluation (an expert review against a list of usability principles), combines 3 factors: how often the problem happens, how hard it is to get past, and whether it keeps coming back. It gives a 0 to 4 scale, from 0 (not a usability problem) to 4 (a catastrophe to fix before release), and says averaging ratings from 3 evaluators works for many practical purposes (NN/g, 1 November 1994, read 6 October 2026). We have borrowed the idea for tester sessions, which is not what that page describes.
Here's where we simplify. After a small test, a 0 to 4 scale can feel like homework, so we sort into 3 buckets:
- Fix before launch: someone could not finish a main task, or finished it wrongly without knowing.
- Fix soon: people slowed down, needed a nudge, or hit a snag that did not block the task.
- Note it: one person, a matter of taste, or a cosmetic snag.
A problem that blocks a main task for even one tester outranks a pile of small annoyances. Our audit report works the same way: it ranks fixes by impact, so what you do first is a decision and not a guess.
Moderated or unmoderated: which fits a small test?
In a moderated test, you (or a friend) watch live, in person or on a call. NN/g's comparison says the facilitator can clarify and probe, and can prompt a quiet tester. In an unmoderated test, the tester does the tasks on their own schedule and records the session for you to review later. The same page says unmoderated sessions can come back incomplete or unusable, and that the format suits specific elements and tight timeframes (NN/g, 12 October 2013, read 6 October 2026).
At small scale, our rule of thumb:
- Start moderated for your first test, because you will see how people think and can fix a muddy task on the spot.
- Use unmoderated to check one fix after you make it, with a tight script and a short task list.
- Do not skip the script in either case. Without one, you are collecting chat.
What did the testers actually do?
Hypothetical example: a library event-signup site, tested with 4 people. The 3 tasks were the ones above, using the story-hour wording. A, B, and C used laptops. D used a phone.
A: Found story hour in under a minute. Signed up without trouble. Spotted a cancel link in the confirmation email and used it.
B: Found story hour quickly. Typed the child's name into the adult field and never noticed. Then hunted the site for a way to cancel, said "I guess you can't," and stopped.
C: Searched "kids" in the site's search box, got no results, and reached story hour through the calendar after 3 minutes. Signed up fine. Could not find a cancel option and said they would phone the library.
D: Scrolled the home page on a phone for a long while. The Events link lived behind a small menu icon D did not recognize as a menu. D gave up on the first task and never reached the other 2.
Now we rate it. Fix before launch: D's dead end on a phone, B's wrong sign-up that nobody caught, and the missing cancel route that B and C could not find. Fix soon: C's empty search for "kids," although the event is for kids, since C still got there. Note it: A wished the footer color were calmer.
That is 5 problems from 4 people, and the ones that matter most (D's dead end and B's unnoticed wrong sign-up) would never have come up in an opinion-only test. Ask these 4 whether they like the site and they might all say yes.
Where does an audit fit?
Our audit clicks through your public pages, so it can flag things like dead buttons (the sample report shows that count) before a tester ever sees them. It is not a security review or a full accessibility audit, and neither is a 4-person test, so if your site handles payments or sensitive data, get proper help for those.
Remember what a test of this size is for: finding problems, not measuring how many visitors will hit them. A script you keep and reuse turns a one-off into a habit, and that habit is what makes your site better one small fix at a time.
