Record As You Go, Not From MemoryDefinitionTest dataThe recorded results of testing a solution: measurements, counts, times, ratings and observations, kept in a form someone else could read and check.Test data is what you write down while the test is happening, and memory is a poor substitute an hour later.Draw up the recording sheet before the first run, so you are filling gaps rather than inventing columns halfway through.Write the raw value you actually observed, such as 3.4 seconds, rather than a tidied version like "about 3".Note the odd things too: the battery was low, the third tester had already seen the app, the floor was carpet instead of vinyl.Date every sheet, because criterion D evidence is stronger when a moderator can see the testing came before the changes.Tables Turn Scribbles Into EvidenceUse one table per test, with the thing you changed down the left and each repeat in its own column.Put the unit in the column heading so the cells hold bare numbers and stay readable.Keep every value to the same number of decimal places, since 3.4 and 3.40 in one column look like two different precisions.Leave a notes column, because "lid popped but the box stayed closed" explains a result that a number cannot.Qualitative results need tables as well: user, task, finished yes or no, time taken, comment.Never delete a result you dislike, and mark anything you exclude as an anomaly with the reason written next to it.ExampleA drop test table has drops 1 to 10 down the side and columns for height in cm, lid shut yes or no, and damage seen.Ten rows of yes and no give you "8 out of 10", which is a real result, unlike "mostly stayed shut".Repeats: One Run Is an AnecdoteA single run can be luck, so repeat each test at least three times under the same conditions.Repeats show the range of results, and a range of 2.1 to 9.8 seconds is a finding by itself.If repeats disagree wildly, something is uncontrolled, so check the tester, the surface and the starting position.For user testing, a repeat means another user, because the same person's second attempt is never a first impression.Five users is a sensible minimum for a school project, and you say so honestly instead of claiming the result speaks for everyone.Averages Give You the Middle, Not the StoryThe mean adds the results and divides by how many there were, which suits times, masses and distances.The mode is the most common answer, and it fits rating scales and yes or no counts better than a mean does.One extreme value drags a mean badly, so 2, 2, 3 and 40 seconds averages to 11.75 seconds even though most users took about 3.Quote the range next to the mean so a reader sees the spread as well as the middle.Round sensibly, since a mean of 3.466666 seconds from a stopwatch you read to 0.1 s should be written as 3.5 s.Percentages need their denominator, so write "4 out of 5 users (80%)" rather than 80% on its own.Common MistakeAn average of two results is a halfway point, not a trend.Turning yes and no answers into an average hides which users failed and why.Charts Earn Their Space Only If They Show SomethingA bar chart compares separate things, like the time each of five users took to finish the sign up task.A line graph only makes sense when the horizontal axis is a continuous scale, such as load in grams against bend in millimetres.Label both axes with the quantity and the unit, or the chart says nothing the table did not already say.Start the scale at zero, because a chopped axis makes a 4% difference look like a landslide.Pie charts suit one set of parts that add up to a whole, such as which of four features users picked as most useful.If the chart just repeats the table without making a pattern easier to see, keep the table and drop the chart.Reading the Data: Hold Each Result Against the ThresholdInterpretation starts by putting each result beside the success threshold you set when you designed the test.Say what the number means in plain words, so 8 of 10 drops held becomes "the lid meets the point in most falls but not all".Look for patterns across tests, such as every failure happening at the same corner joint.Keep what the data shows separate from what you think caused it, and label the cause as your explanation.Watch for results that answer a different question, like a fast load time measured on a page the browser had already cached.State how confident you are, since five users and three repeats supports a modest claim rather than a certainty.Exam techniquePut dated photographs of the test in progress next to the table, so a moderator can see the numbers came from real runs.Show the raw tables in your ePortfolio, not only a tidy summary chart.A Failed Test Is Still a Result Worth HavingData showing the solution missed the specification is evidence, and writing it earns credit.Record how far short it fell, since 9.2 seconds against a 6 second target is a different problem from 45 seconds.Ask whether the fault sits in the product or in the test, because a snapped clamp may have been loaded with twice the mass the client would ever use.Failures point straight at your improvements, so a clear failure is often the most useful data you collect.Active recallWhy does a results table need the unit in the heading and a notes column?Give one reason to repeat a test three times, and say what a wide range between repeats tells you.When is the mode a better summary than the mean?What is wrong with a bar chart whose vertical axis starts at 80 instead of 0?Your solution failed a test badly. Name two things you should write down about that failure.