Usability Testing Means Watching, Not AskingA usability test hands one real user a task on your design and records what actually happens while you say as little as possible.Asking "do you like it" gets you politeness, while setting a task gets you behaviour you can count.Testers should match your persona, so a product aimed at year 7 students cannot be tested only on your design teacher and your best friend.Three tasks in about fifteen minutes is a sensible session, since eight tasks exhausts everybody and blurs the results.Test whatever exists, because paper, a clickable prototype and a finished build all surface real problems.Write Tasks, Not QuestionsA task is a goal with no instructions hidden inside it, such as finding out whether the library has a particular book and when it comes back.Never name a button in the task, because "tap Search and type the title" tests whether they can follow orders.Give the task a reason so it feels real, such as needing the book for tomorrow's lesson.Write the success condition before the session, so you are not inventing the pass mark once you have seen the result.Order the tasks so an early one does not teach the answer to a later one.ExampleWeak task: "Use the search bar to look up a book."Better task: "You need a copy of The Hunger Games by Friday. Find out whether you can get one."Success condition: the tester says the loan status out loud within 60 seconds, unprompted.Think Aloud Turns Silence Into DataIn the think aloud protocol you ask the tester to keep saying what they are looking for, expecting and thinking as they go.It exposes the reasoning behind a wrong tap, for instance "I thought the star meant save it, not rate it".People go silent when they concentrate, so prompt with "what are you thinking now" rather than with a hint.Answer no questions during the task, and use the standard reply: "what would you do if I was not sitting here".Save your own questions for a debrief at the end, when asking them can no longer contaminate the test.Common MistakeThe strongest urge in a first test session is to help, and helping destroys the data you came for.A nod, a glance towards the right button or a quiet "almost" is enough to steer a tester to the answer.Sit slightly behind them, keep your hands still, and write rather than talk.Measure Four Things Every SessionTask completion is the headline result, recorded as finished, finished with difficulty, or gave up.Time on task runs on a stopwatch from the moment you stop reading the task out to the moment they succeed or quit.Errors are wrong taps, wrong paths and text typed in the wrong box, counted as a tally rather than judged.Hesitations are pauses longer than about three seconds, and where they land is usually the most useful thing you take away.Note the first click as well, because the first tap reveals what the tester thought the screen was offering.Ask for a 1 to 5 difficulty rating straight after each task, which gives you one comparable number per task per person.Five Users Find Most of the ProblemsTesting with about five users uncovers the large majority of usability problems in a design.The reason is overlap, since the sixth and seventh testers keep walking into the same walls the first five already found.Three rounds of five beats one round of fifteen, because the rounds give you space to fix things in between.It is five per user group, so a product with two genuinely different personas needs two sets of five.Five people is enough to find problems and nowhere near enough to prove a preference, so never report percentages off five testers.Exam techniqueCriterion D asks you to test against your own design specification, so map each task to one specification line.Put raw session notes in the appendix and a summary table in the main body of the folder.Record the date, the prototype version and which user group each tester belonged to next to every result.Findings Are Worthless Until Something ChangesWrite the session up within an hour, because the detail you did not note down disappears quickly.Combine findings from all testers into one list and put a number next to each problem for how many people hit it.Sort by severity: blocks the task, slows the task, or merely irritates.Fix the blockers first even though the small irritations are easier and far more tempting to tick off.Write each change in three parts, finding, change and reason, and keep the before and after screenshots side by side.Retest the changed screens with fresh testers, since anyone who struggled the first time now knows where the button is.Active recallWhy does setting a task work better than asking a question?What is the think aloud protocol, and what do you say when a tester goes silent?Name four things worth measuring during a session.Why is three rounds of five testers better than one round of fifteen?How do you decide which problem to fix first?