A reaction time study is one of the few school experiments that produces genuine, unfaked data in a single lesson, needs almost no equipment, and gives every student a personal stake in the result. This page is a complete template: a hypothesis you can adapt, a variables table, a step-by-step method, a blank data table, a worked graph, and an honest list of the errors that will show up in your results whether you mention them or not.
Everything here assumes a class-sized study - around ten participants, five scored trials each. Scale it up if you have time, but do not scale the trials down: fewer than five per person and your averages will be measuring luck.
Pick exactly one thing to change. The most reliable choices for a school setting are stimulus type (visual versus auditory), task type (simple versus choice), hand used (dominant versus non-dominant), or condition (before versus after a distraction, exercise, or a set number of practice trials).
Write a directional hypothesis - one that predicts which way the result will go, so that the data can actually disagree with you:
Avoid the classic weak version - reaction time will be affected - which cannot be wrong and therefore cannot be tested.
| Type | In this experiment | How it is handled |
|---|---|---|
| Independent | Stimulus type (visual / auditory) | The one thing you deliberately change |
| Dependent | Reaction time in milliseconds | Measured, five trials per condition |
| Controlled | Device and browser | Same machine for every participant |
| Controlled | Hand used | Dominant hand only, index finger |
| Controlled | Wait before signal | Randomised 1.5-4 s every trial |
| Controlled | Practice | Two unscored trials for everyone |
| Controlled | Environment | Same room, notifications off, no spectators calling out |
| Confounding | Age, fatigue, caffeine | Record them; discuss them in the evaluation |
If you are using a screen-based test, our three-round reaction test covers the simple and choice conditions in one sitting and marks early taps as faults automatically. For a screen-free version, the ruler drop method needs only a metre rule and converts catch distance directly to milliseconds.
Copy this table into your book, one row per participant per condition. Record every trial - editing out the ugly ones is falsifying data, and the spread is itself a finding.
| ID | Age | Cond. | T1 | T2 | T3 | T4 | T5 | Mean | Range | False starts |
|---|---|---|---|---|---|---|---|---|---|---|
| P1 | 15 | Visual | 281 | 264 | 310 | 272 | 258 | 277 | 52 | 1 |
| P1 | 15 | Auditory | 246 | 238 | 259 | 231 | 240 | 243 | 28 | 0 |
| P2 | 15 | Visual | ||||||||
| P2 | 15 | Auditory | ||||||||
| ... |
P1 is filled in as a worked example. Note the range column: that participant's visual trials varied by 52ms while their auditory trials varied by 28ms, which is worth a sentence in the analysis on its own.
| Statistic | How to calculate | Why it matters |
|---|---|---|
| Mean | Sum of five trials ÷ 5 | Comparable with published figures |
| Median | Middle value of the five | Ignores one lapse of attention |
| Range | Slowest − fastest | Shows consistency, not just speed |
| Group mean | Mean of all participant means | The headline number per condition |
| Difference | Group mean A − group mean B | Tests the hypothesis directly |
Report the mean and the median together. Reaction time data is right-skewed - a single distracted trial can drag a mean up by 20ms while barely moving the median - so quoting both shows you understand your own data rather than just averaging it.
A bar chart of group means, one bar per condition, is the standard presentation. Put the condition on the x-axis, reaction time in milliseconds on the y-axis, start the y-axis at zero, and label both axes with units. Here is what a finished chart looks like for a three-condition version of this study.
Two upgrades that earn marks: add error bars showing the range or standard deviation for each condition, and include a scatter plot of individual means against age if your participants span more than a couple of years. If you are testing the same people twice, a paired line graph - one thin line per participant, condition A to condition B - shows instantly whether the effect held for everyone or only for a few.
| Source | Effect on results | Control |
|---|---|---|
| Anticipation | Artificially fast trials | Randomise the wait; discard early responses |
| Practice effect | Later condition looks faster | Alternate the order between participants |
| Device latency | All scores inflated 20-70ms | One device throughout; do not compare with other classes |
| Fatigue | Later trials slower | Rest between conditions; keep sessions short |
| Distraction | Random slow outliers | Quiet room; report the median as well as the mean |
| Small sample | Difference may be chance | Ten participants minimum; state the limitation |
In the evaluation, resist the urge to claim more than you measured. A 30ms difference between two conditions with ten participants is suggestive, not proven, and saying so is a stronger conclusion than overclaiming. Compare your group means against the published ranges on the average reaction time by age page and, for teenage participants, the 15-year-old chart. Use the conversion tables to turn ruler catches into milliseconds, and read simple vs choice reaction time before you compare any two task types.
The conclusion is one or two sentences and nothing more: state the group means, state the difference, and say plainly whether the data supports your hypothesis. For example - the mean reaction time to an auditory signal was 235ms compared with 273ms for a visual signal, a difference of 38ms, which supports the hypothesis that auditory signals are responded to more quickly.
The evaluation is where the marks live. Say how confident you are and why: how many participants, how consistent the individual results were, whether any participant showed the opposite pattern, and which of the errors above you think mattered most. Then propose one specific improvement that follows from your own data rather than a generic one - if your ranges were wide, propose more trials per participant; if your two conditions were run in the same order for everyone, propose full counterbalancing; if scores clustered oddly, propose checking device latency with a second machine.