Reaction time varies from attempt to attempt. A single result could be unusually fast because you anticipated the signal, or unusually slow because attention drifted.
With repeated trials, those one-off events have less power to define the whole session.
Repeated-testing research shows that participants often improve across early sessions or trials. Part of this is simply becoming familiar with what the test expects.
A short warm-up helps separate “learning the interface” from the set you want to compare later.
There is no universal magic trial count. An older reliability study found that, under its particular procedures, about 18 trials for a simple reaction task and 30 trials for a two-choice task produced high reliability. Different equipment, populations and statistics can require different amounts.
Save the median or average of the recorded trials, plus the best and any false starts. If you are comparing conditions, use the same summary metric every time.
For Aim, Choice and Fakeout modes, also retain accuracy or error information because latency alone does not capture the entire task.
Five is much better than one for casual testing, although a larger set gives a more stable estimate.
If your goal is a stable baseline, it is often useful to practise first and record the later block separately.
More trials reduce measurement noise and help researchers estimate smaller effects more reliably.