Prototype User Testing: Plan Tests That Improve the Design

TL;DR
Run a focused prototype test with 4 to 8 actual or likely users, give each person believable tasks without naming the interface, and watch where they hesitate, fail, recover, or misunderstand. Use the smallest prototype that answers the decision at stake, then fix repeated task blockers before development.
Run a focused prototype test with 4 to 8 actual or likely users, each attempting believable tasks without instructions. Choose only enough prototype detail to answer the design decision, then change the design where people repeatedly hesitate, fail, or misunderstand what happens next.
A prototype test is not a design review. It is an observation session: users try to achieve a goal while you learn whether the flow, language, and feedback help them succeed.
Choose the Test That Answers Your Decision
Start with the decision that development cannot safely make on assumption. A paper sketch can answer whether a concept makes sense. A realistic clickable flow can answer whether people can complete a payment, application, booking, or account task.
Use this decision table before you recruit anyone.
| Decision | Prototype and Participant | Task Prompt | Method |
|---|---|---|---|
| Concept understood? | Paper storyboard; recent problem experience | Explain how you solve that problem today. | Moderated interview |
| Route easy to find? | Clickable wireframes; likely users | Find a way to pay a bill. | Moderated usability test |
| Core task completable? | Clickable flow; actual or likely users | Buy running shoes for delivery this week. | Moderated usability test |
| Wording and states clear? | Realistic screens; affected users | Change a delivery address after checkout. | Moderated usability test |
| Stable alternatives compared? | Near-final flows; screened users | Complete one defined task independently. | Unmoderated task test |
The table means you should not build a polished prototype to answer an early concept question. It also means you should not use a rough storyboard to judge detailed interaction, confirmation messages, error recovery, or trust.
Digital.gov guidance distinguishes static prototypes for intent and design feedback from functional prototypes for observing interaction. GOV.UK guidance makes the same point: early questions can use simple, low-tech prototypes, while later questions need a design closer to the intended product.
Decide What Evidence You Need
Write one decision statement before making the prototype. A good statement names the user, the task, and the design choice that will change.
Use this format:
We need to decide whether first-time customers can complete checkout without help.
We will watch customers add an item, choose delivery, and pay.
We will redesign any step that blocks two or more participants from continuing unaided.
That statement stops a session becoming a request for general feedback.
Genuine usability evidence is observable:
- A participant chooses the wrong route.
- A participant cannot explain a label.
- A participant pauses because the next action is unclear.
- A participant makes an error and cannot recover.
- A participant expects a different result after tapping a control.
Polite feedback is useful context, but it is weaker evidence:
- “I like the design.”
- “This looks clean.”
- “I would probably use it.”
- “You should add this feature.”
Ask for opinions after the task. Base design priorities on what users tried to do before they heard your explanation.
A usability test observes people attempting to use a product while thinking aloud. Digital.gov’s method recommends agreeing the scenarios, users, moderator, and observers before the session.
Recruit People Who Actually Face the Problem
Recruit participants because they use the product category, face the underlying problem, or have the role your product supports. Do not recruit friends, colleagues, or designers unless they genuinely match that profile.
For a grocery-delivery checkout, recruit people who order groceries online. For a B2B approval flow, recruit people who approve invoices or manage the relevant workflow. For an accessibility feature, include people who use the relevant assistive technology in daily life.
Our recruiting rule is simple: ask about a recent real behaviour, not a future intention. “Have you ordered groceries online in the past month?” is stronger than “Would you use a grocery app?” GOV.UK recommends recruiting people who may have been in the relevant situation within the past 6 months, so their responses come from experience rather than imagination. Recruitment guidance
Plan a small qualitative round first. Four to eight participants can reveal recurring barriers in a defined flow, then a second round can test whether the redesign solved them. A small round does not estimate the percentage of all users who will fail. Use a larger, purpose-built quantitative study if you need a benchmark or population-level result.
Write Prototype User Testing Questions That Do Not Lead
Write tasks as situations and goals. Do not name the navigation item, feature, or wording you want the participant to find.
Use this task structure:
You need to pay a bill that is due today. Show me what you would do.
That prompt lets you observe the participant’s natural route. Compare it with a leading prompt:
Use the Payments tab to pay a bill due today.
The second prompt tests whether the participant can follow your instructions. The first tests whether the interface makes sense.
Use these questions during or after a task:
- “What are you looking for here?”
- “What do you expect to happen?”
- “What made you choose that?”
- “What would you do next if I were not here?”
- “What did you expect to find on that screen?”
- “How would you describe this option in your own words?”
Avoid asking “Was that easy?” while the participant is still trying. The question can make a person defend the design or rush to finish.
Good tasks have a clear goal, feel believable, and do not reveal the answer. Task-design guidance also recommends using a discussion guide so every participant receives consistent instructions.
Run a 45-Minute Moderated Session
Our starting session plan uses 45 minutes, three core tasks, one moderator, and one note-taker. GOV.UK says moderated usability sessions commonly run for 30 to 60 minutes. Session guidance
Use this script.
-
Set expectations, 3 minutes. Say: “We are testing the design, not you. Some parts are unfinished. Please say what you are thinking as you go.”
-
Learn the participant’s context, 5 minutes. Ask how they solve the problem today. This gives you language and habits to compare with the prototype.
-
Run the three tasks, 27 minutes. Read each scenario once. Stay quiet while the participant works. If the person gets stuck, ask what they expected to happen before offering help.
-
Probe observed moments, 7 minutes. Ask about hesitations, wrong turns, and successful workarounds. Ask what the person expected at each important point.
-
Close, 3 minutes. Ask for final thoughts only after the tasks. Thank the participant and explain what happens to the research notes or recording.
Record the screen and audio only with informed consent. If a session uses participant details, documents, or recordings, plan the data handling before recruitment. Our privacy and anonymisation information explains how Qualfacto handles participant data.
Choose Moderated or Unmoderated Testing
Choose moderated testing when you need to understand why a person stalled, mistrusted a screen, or chose an unexpected path. Moderation also helps when the prototype has dead ends, the task is sensitive, the user needs support, or the team must test unfamiliar interactions.
Use an unmoderated task test when the prototype is stable, the task is short, and the completion point is obvious. An unmoderated test can show repeated patterns across more completed sessions. It cannot reliably explain why a person became confused.
Our rule: use moderated testing to discover the problem, then use unmoderated testing to check a stable, self-contained task. Remote sessions can make it harder to guide a participant and understand exactly how they are using a prototype. Remote-testing guidance
A prototype session also cannot validate everything. Test technical performance, security, real payment processing, and live data separately. If real-world context affects behaviour, add contextual research rather than forcing that question into a prototype session. Digital ethnography can reveal the habits and environments that a clickable flow cannot reproduce.
Turn Sessions into Design Priorities
Take notes on behaviour, not interpretations. Write “looked for delivery price in basket for 42 seconds” instead of “delivery price was confusing.”
After each session, capture three findings:
Observed behaviour: The participant searched the basket for the delivery price.
Impact: The participant paused before checkout and questioned the total.
Design change: Show delivery cost before the payment step.
At the end of the round, group similar observations. A repeated problem is one finding with several pieces of evidence, not several unrelated complaints.
Prioritise work in this order:
- Task blockers: Users cannot complete a core action.
- Repeated confusion: Several users misunderstand the same label, control, or state.
- High-cost errors: A mistake could lose money, expose data, create rework, or reduce trust.
- Friction: Users complete the task but take an indirect or effortful route.
- Suggestions: Users request an improvement without an observed breakdown.
Then change the prototype and test the changed part with a new round. Iteration matters because a fix can create a new misunderstanding elsewhere in the flow.
FAQs
Should I Show Customers an Unfinished Prototype?
Yes, tell participants that the prototype is unfinished and that some controls may not work. That framing encourages criticism and prevents participants from treating missing functionality as a personal failure. It also gives you permission to ask what they expected at a dead end.
Can I Test a Prototype with Assistive Technology Users?
Yes, and participants should use their own device and assistive technology when practical. Personal configurations are difficult to recreate in a research lab or shared screen. Focus on the real task the person needs to complete, not a generic accessibility checklist.
Should Observers Speak During the Session?
No, observers should take notes and leave moderation to one person. Multiple voices can lead the participant, interrupt their thinking, or turn the session into a stakeholder review. Collect observer questions after the task or at the end of the session.
What If the Prototype Cannot Show a Necessary Screen?
Mark the missing interaction in advance and decide how the moderator will handle it. Ask what the participant expected to happen before moving them to the next prepared screen. Treat repeated expectations as evidence of a missing path, label, or system state.
At Qualfacto, we manually review and match panel members to UX tests based on their career, lifestyle, and background. Explore our qualitative research panel when you need participants for a prototype-testing round.



