● Daily

Practical guide

How to Measure Your Progress in Brain Training

By LaserMind Team ·

Why a rising score is hard to read

When a score improves, three things may have changed: you, your familiarity with the test, or the conditions. Only the first is the progress you are after.

You get better at taking the test. A meta-analysis of 50 studies (107 samples, 134,436 people) of cognitive ability tests taken more than once found an adjusted average retest gain of 0.26, a small effect; gains were larger when people were coached and when the identical form of the test was used again [1]. On a game like n-back it is easy to see where some of that comes from: by the third session you know the keys, the rhythm and what a match looks like.

Untrained groups improve too. The same happens in training research. In a randomised study of 102 healthy young adults, a group that practised the Stroop task for three weeks, an active control group and a group with no training all improved substantially from the first to the second testing on every untrained task [2]. Only the comparison with the control groups showed that the Stroop training itself had added no transfer [2].

The game score is not your everyday ability. A major review of brain-training studies found extensive evidence that training improves the trained tasks, less evidence for closely related tasks, and little evidence that it improves everyday cognitive performance [3]. A new n-back level is real progress at n-back. Whether you also follow a meeting better is a separate question that needs a separate measure.

Conditions and chance vary. Response times depend on your device and browser, so a phone session and a laptop session are not directly comparable. And an unusually good session tends to be followed by a more ordinary one, and a bad one by a better one; statisticians call this regression to the mean. A single score, high or low, says little.

How LaserMind measures progress

Each part of your LaserMind results is designed around these problems.

One main number per exercise, and a personal best per mode

ExerciseMain numberBetter is
N-backThe highest N you played with at least 80% correctHigher
StroopInterference cost: median time on mismatching words minus matching ones, correct answers onlyLower
Schulte tableTime to complete the gridLower
Number spanLongest sequence you got right at least onceHigher
Go/No-GoAccuracy; your personal best is kept by d′ (see below)Higher
Switch the ruleSwitch cost: median extra time just after a rule change, correct answers onlyLower

Personal bests are kept separately for each mode: single and dual n-back, each Stroop mode and language, each Schulte grid size. As a guest, your history stays on your device; with an account, bests are also kept per device type (desktop, tablet, phone).

Medians rather than averages

The Stroop interference cost, the switch cost, Find the Odd One's search time and Mental Rotation's response time all use the median of your correct answers: the middle value. One slow response, because you sneezed or glanced away, barely moves a median but can drag an average. Answers faster than 150 milliseconds count as errors, because they are too fast to be real responses.

d′: accuracy you can't inflate by pressing more

In n-back and Go/No-Go you could catch every target by pressing all the time, at the cost of many false presses. d′ (d-prime) combines hits and false presses into one number that shows how well you tell targets from non-targets, regardless of how often you press. From 3-back, near-misses (a match one step too early or too late) are mixed in, so guessing shows up as false presses. If your N stays the same but d′ rises, you are separating matches more cleanly.

A short trend after every session

The results screen shows your recent sessions in the same mode, up to six from oldest to today, and says how steady they are: "steady" when their spread is less than about 10% of their average, "varies a lot" when it is more than about 25%. The "What happened" panel shows where your errors came from, such as missed targets, presses you should have held back, or mistakes just after a rule change. Working on one kind of error at a time is the idea behind deliberate practice.

The three-level report

Your report asks three separate questions:

  1. Trained tasks. For each exercise and mode with at least four full sessions, it compares the median of your first three sessions with the median of your last three (two each when you have only four or five). Changes within 3% either way are shown as "about the same". Daily Circuit blocks are left out because they are shorter versions, and so are Kids-mode sessions.
  2. Untrained tasks. The LaserMind Check-in is four short tasks, about ten minutes in all, in versions that our programmes and the Daily Circuit don't train: arrows (answer the direction, not the side), the light path backwards, rule switching with many switches, and digit span backwards. You take it once as a starting point and then about every four weeks. Each check-in uses a new version of the same difficulty.
  3. Real life. Focused minutes and distractions per hour from the Deep Work Timer, and your answers in the weekly check-in of the Reflection tool: how much you still remember of what you learned the week before and, if you choose, how much you slept.

The report says plainly what to expect. Level 1 almost always improves. Little or no change at level 2 is the most common pattern, which fits what reviews of training studies find [3]. And because there is no control group, this is self-tracking, not a study.

Step by step: reading a trend over noise

These steps are practical advice, not research findings.

  1. Fix the comparison. Choose one exercise, one mode and one device, and stick with them. The report groups sessions by exercise and mode, so if you switch between devices, its trend can mix them.
  2. Keep the conditions similar. Play at a similar time of day, in a similar setting, with the same input (keyboard or touch). For Listen & Count, use the same headphones at the same volume.
  3. Let the learning phase pass. Expect the first few sessions to rise quickly while you learn the task, and judge the trend after that.
  4. Compare groups of sessions, not single ones. A personal best is your best day, not your usual level; the median of three sessions is steadier.
  5. Check the spread before the change. If recent sessions "vary a lot", a change of a few per cent is within the noise. Look for a shift larger than your usual variation that holds over several sessions.
  6. Keep the check-in tasks untrained. The Arrows mode of the Stroop game, backward Light Path and Number Span, and rule switching at 50% are all available in normal training. If you practise them, the check-in no longer measures transfer for you. LaserMind suggests about four weeks between check-ins, so that each one depends less on remembering the last.
  7. Measure the goal itself. If your goal is focus at work, the Deep Work Timer's focused minutes are closer to it than any game score.

Common mistakes

  • Mistaking early gains for transfer. Early jumps on the trained game are partly just learning the task.
  • Reading one check-in as a verdict. A single change in either direction is weak evidence. Look at several check-ins.
  • Reading the distraction count on its own. The Deep Work Timer counts the distractions you notice. As you get better at noticing, the count can rise for a while, so read it alongside your focused minutes.

What the research says

◆ SupportedScores rise when the same test is repeated, even without training [1]. Practice improves performance on the trained task [3].
◇ ConditionalGains on closely related tasks: reliable in the short term after working-memory training, but verbal working-memory gains were not sustained at follow-up [4].
✕ Not establishedThat a higher score on a practised game means better everyday memory or attention [3].

Try it on LaserMind

▶ N-Back

Play a few sessions on the same device and watch both your N and your d′.

Try it on LaserMind

▶ Deep Work Timer

Track the outcome closest to real life: minutes of focused work on one real task.

What research has not established

  • The retest meta-analysis pooled studies of standard cognitive ability tests, not game-like exercises repeated dozens of times. How large practice effects are on LaserMind's tasks is not known.
  • LaserMind's report has no control group, so it cannot separate practice effects from real change. A rise at level 2 may still be practice.
  • New versions of the check-in tasks reduce, but cannot remove, the advantage of having done a task before: you still know the rules.
  • Self-ratings in the weekly check-in are open to expectations, a problem that even formal trials rarely control [5].
  • The thresholds in the report (four sessions, 3%, 10% and 25%) are practical rules chosen by LaserMind, not research standards.

Frequently asked questions

Does LaserMind compare me with other people? No. Scores are compared only with your own history. We never rank you against "the population" or convert scores into IQ.

My score dropped. Did I get worse? Probably not. Single sessions vary, and a drop after a personal best is often just a return to your usual level. LaserMind exercises are practice, not diagnostic tests. If you notice real, lasting changes in memory or concentration in daily life, talk to a doctor or another qualified professional.

References

  1. Hausknecht, J. P., Halpert, J. A., Di Paolo, N. T., & Moriarty Gerrard, M. O. (2007). Retesting in selection: A meta-analysis of coaching and practice effects for tests of cognitive ability. Journal of Applied Psychology, 92(2), 373–385. https://doi.org/10.1037/0021-9010.92.2.373
  2. Talanow, T., & Ettinger, U. (2018). Effects of task repetition but no transfer of inhibitory control training in healthy adults. Acta Psychologica, 187, 37–53. https://doi.org/10.1016/j.actpsy.2018.04.016
  3. Simons, D. J., Boot, W. R., Charness, N., Gathercole, S. E., Chabris, C. F., Hambrick, D. Z., & Stine-Morrow, E. A. L. (2016). Do "brain-training" programs work? Psychological Science in the Public Interest, 17(3), 103–186. https://doi.org/10.1177/1529100616661983
  4. Melby-Lervåg, M., & Hulme, C. (2013). Is working memory training effective? A meta-analytic review. Developmental Psychology, 49(2), 270–291. https://doi.org/10.1037/a0028228
  5. Boot, W. R., Simons, D. J., Stothart, C., & Stutts, C. (2013). The Pervasive Problem With Placebos in Psychology: Why Active Control Groups Are Not Sufficient to Rule Out Placebo Effects. Perspectives on Psychological Science, 8(4), 445–454. https://doi.org/10.1177/1745691613491271