Cognitive control
Stroop test with your interference in milliseconds
Name the ink a color word is printed in, ignore the word, and the run reports how much slower the mismatched trials were than the matched ones — 72 scored trials in about four minutes, free, with no account and nothing held back. A third of the trials show XXXX instead of a word, which is what lets the result separate the help a matching word gives you from the cost a mismatched one imposes. If you cannot tell the four inks apart, the counting variant on the same page measures the same conflict with no hue in it.
- 100% free
- No signup
- 72 scored trials
- Counting variant, no color
- Interference in ms
72 scored trials in three passes of 24, about four minutes. The number this produces is a difference between two of your own conditions, so your monitor, your keyboard and how awake you are cancel out of it — which is the only reason a reaction time measured through a browser is worth printing at all.
The color-word run, in full
- Stimulus
- one of RED, BLUE, GREEN, YELLOW or XXXX, printed in one of four inks
- Your job
- press the key for the ink, whatever the letters spell
- Keys
- D = red, F = blue, J = green, K = yellow — fixed, so two runs are comparable
- Trials
- 72 scored, 24 in each of the three cells, six per ink per cell
- Timing
- 400 ms cross, then a gap drawn between 300 and 700 ms, then the word until you answer or 2 s passes
- Void
- a trial answered inside 150 ms, one answered before the word appeared, and any trial the tab was hidden during
The practice trials have no time limit at all and tell you after each one whether the key was right. Nothing in them reaches the result.
How to take the Stroop test
Four keys, three kinds of trial, and one difference at the end of it.
Pick the color task or the counting task
The color task shows a word like BLUE printed in some ink and asks for the ink: D for red, F for blue, J for green, K for yellow, with the same four as buttons if you have no keyboard. The counting task stacks one word one to four times in a single gray and asks how many lines there are. Take the counting one if telling red from green is work for you — it is the same conflict without the hue, and it is offered rather than assigned because a page that guesses at your eyesight will guess wrong.
Run 12 practice trials with no clock on them
Practice trials have no response window at all and tell you after each one whether the key was right, so the mapping is in your fingers before anything counts. None of them reaches the result. Skip them if you have done this before; the scored block reruns nothing you have already learned.
Answer 72 trials, resting twice, then read the difference
The scored block runs 24 trials, rests, runs 24 more, rests, and finishes. Keep both index fingers on the keys and answer as fast as you can be right. The result leads with incongruent minus congruent in milliseconds and a 95% interval on that one difference, then splits it into facilitation and cost, then shows each cell's own median and accuracy so you can see whether speed was bought with errors.
Technical specifications
| Trials | 72 scored — 24 congruent, 24 neutral, 24 incongruent, six per ink in each cell — plus 12 optional practice trials that are excluded from every figure |
|---|---|
| The three cells | Congruent is BLUE in blue ink; incongruent is BLUE in red ink; neutral is XXXX in blue ink. The neutral string is the modern control — Stroop's own control in 1935 was a solid color square, which has a different visual form from a word |
| Response | Four keys read by physical position, so D, F, J and K are the same four keys on a QWERTY, AZERTY or Dvorak layout. The counting variant takes 1 to 4 on the number row or the keypad. Four on-screen buttons work throughout |
| Trial structure | 400 ms fixation cross, then a blank gap drawn uniformly between 300 and 700 ms, then the stimulus until you answer or 2,000 ms passes, then 400 ms blank. The gap is jittered because a fixed one is learnable — after a dozen trials you would be timing the rhythm rather than reading the screen |
| What voids a trial | A response inside 150 ms of the stimulus painting, a response that arrives before it paints at all, and any trial during which this tab went to the background. An interrupted trial is added back at the end so the run still delivers 72 |
| Timing floor | Your display's refresh rate is measured while the start screen is up and the resulting frame time is printed with the result. On a 60 Hz panel one frame is about 16.7 ms, and that much of every figure is display and input hardware before any of it is you. Differences are reported to the whole millisecond for that reason |
| Reference figure | 50-200 ms — MacLeod (1991), Half a century of research on the Stroop effect: an integrative review, Psychological Bulletin. It is a span across published studies, not a distribution, so this page compares your difference with it and never converts it into a rank |
| What leaves the page | Nothing, unless you press the copy button, which writes plain text to your own clipboard. There is no request to any server during a run, and no trial-by-trial times are in the copied summary |
Frequently asked questions
Why is my interference effect smaller than the numbers I have read about?
Because you answered with a key rather than with your voice, and that alone roughly halves the effect. The large figures quoted from the literature come from experiments where participants said the color out loud, where the word and the response share a channel — the mouth is already loaded with the wrong word before the ink has been named. A key press does not have that problem, so a keyboard version of the task sits at the low end of the published span by construction rather than by anybody underperforming.
Why does a third of the run show XXXX instead of a word?
Because incongruent minus congruent is two effects added together and this is how they come apart. A word that names the ink helps you, a word that names a different ink hurts you, and the single difference everyone quotes contains both without saying how much of it is which. The XXXX string names no color, so neutral minus congruent is the help and incongruent minus neutral is the harm. Most published work finds the harm is the larger of the two, and a run where yours is the other way round is worth a second run before it is worth a theory.
I am color-blind. Is there any point in me taking this?
Take the counting variant instead — it is on the start screen, not behind anything, and it needs no hue at all. The color task would measure how well you separate red from green, dressed up as a measurement of cognitive control, and the number it returned would be about your retina. The counting variant stacks one word one to four times in a single gray and asks how many lines there are while the word says a different number, which is the same conflict between a name and a property. The two scores are not on one scale and this page never adds them.
Does a large effect mean I have an attention problem?
No, and the reason is arithmetic rather than caution. The color-word effect is one of the most reliably reproduced findings in psychology at the level of a group, and a poor separator of one person from another: measure the same person twice and their two interference scores agree much less well than the size of the group effect suggests. That gap between a robust phenomenon and an unstable individual measure is documented directly, and it means a single browser run cannot place you against anybody. Nothing on this page is a diagnostic instrument and no result from it should be shown to a clinician as though it were data.
Can I compare this number with one from a lab or a clinic?
Not directly, for four reasons that all point the same way. The response is manual here and usually vocal there; the trials are mixed here and often blocked on a printed card there; the display and keyboard add a floor of roughly one frame that a laboratory measures and removes; and there is nobody sitting beside you correcting an error the moment it happens. What does transfer is a comparison of you with you — the same variant, on the same machine, a week apart.
What happens if I hit a key before the word appears?
The trial is voided and the run carries on. There is no stimulus onset to measure that press from, so any number produced would be an interval between a fixation cross and a finger, which is not what this page reports. The same rule catches responses that arrive within 150 ms of the word appearing: light has to reach the retina and a signal has to reach the finger before any decision is possible, so a figure below that floor is a key already on its way down. Both are counted and named in the result rather than quietly averaged in.
Why are there rests after trial 24 and trial 48?
Because 72 unbroken trials would measure stamina alongside conflict, and the run is meant to measure only one of those. Interference is sensitive to how fresh the eye and the hand are, so a block long enough to tire you produces a difference that is partly a fatigue curve. Two rests you control also mirror the shape of the original card task, which was three separate passes rather than one long one. Nothing is timed while you are resting.
What Stroop measured in 1935, and what this version changed
The 1935 experiments did not contain the comparison everyone now calls the Stroop effect. Stroop compared naming the ink of conflicting color words against naming the color of solid squares, and reading conflicting color words against reading the same words in black — and the striking half of the result was the asymmetry. Reading was barely touched by the ink; naming was slowed enormously by the word. That is the finding: one of the two operations is automatic enough to run through interference and the other is not. The congruent condition, where the word names its own ink, arrived later, and with it the difference this page leads with. Anybody quoting a figure for “the Stroop effect” is quoting one of several different subtractions, which is a large part of why the published numbers span the range they do: 50-200 ms, per MacLeod (1991), Half a century of research on the Stroop effect: an integrative review, Psychological Bulletin.
Three things move that figure around, and a page that does not name them is not describing its own task. The first is the response channel — saying the color out loud produces far more interference than pressing a key, because speech and the printed word compete for the same output, and a keyboard version therefore sits low in the range on purpose. The second is whether the conditions are blocked or mixed: a printed card of nothing but incongruent items lets you settle into a strategy, while a shuffled run denies you that, and the two designs are not measuring quite the same thing. The third is the control condition, and it is the one browser versions usually drop. Without a neutral cell there is no way to say how much of a difference was a mismatched word hurting and how much was a matching word helping, so a two-condition test reports a quantity it cannot decompose and calls it interference. This run keeps the third cell.
The reason no rank or percentile appears beside your milliseconds is specific and worth stating: the effect is robust as a phenomenon and unreliable as an individual measure. The same person tested twice produces two interference scores that agree far less well than the size of the group effect implies, which is exactly the property that makes a task good for demonstrating a mechanism and bad for sorting people — Hedge, Powell & Sumner (2018), The reliability paradox: Why robust cognitive tasks do not produce reliable individual differences, Behavior Research Methods sets this out directly. So the page hands you a difference and a confidence interval on it, and leaves the ranking alone. If you want to see the same control mechanism measured through a different conflict, the neighbors are worth a run each: the flanker task puts the distraction beside the target rather than inside it, and the Simon task uses a distractor with no relevance to the answer at all — only to the hand that gives it. A broader sweep across several of these paradigms at once is what the executive function assessment assembles, and the selective attention test asks the filtering question without a speeded response at all.
Why a second route through this page exists
Red-green color vision deficiency, prevalence: ≈ 8% among men of European descent, against roughly 0.4% among women.
Birch (2012), Worldwide prevalence of red-green color deficiency, Journal of the Optical Society of America A
The 8% figure is specific to men of European descent and is lower in several other populations, so it describes who this page has to keep a second route open for rather than the odds for any particular reader. Nothing on this page attempts to detect the condition; the alternative task is offered to everyone, up front.
A reaction time here is the interval between the frame that painted the stimulus and the timestamp the browser attached to your key, both read from the same monotonic clock. What neither can see is the display pipeline behind it, so on a 60 Hz screen roughly 16 ms of every figure below is the machine rather than you. That is the timing floor: two numbers closer together than that are the same number, and this page reports no precision it cannot support.
This is a measurement exercise, not a clinical assessment. It reports what you did on this page against a stated reference and nothing more — it cannot establish an attention disorder, a reading difficulty or a frontal-lobe problem. Only a qualified professional, working with more than a browser, can make that judgment.
Where your 72 response times are worked out
Every number on this page is worked out by JavaScript running in the tab you are reading it in. Your answers, your reaction times and your score are never uploaded, logged or kept — which is also why the test carries on working after you disconnect from the network, and why nothing here can be held back behind an email address.
Every keypress timestamp lives in a JavaScript array inside this tab and is used once, to subtract one cell mean from another. The array is not written to storage, not sent anywhere, and gone when the tab closes — which is also why reloading the page loses a finished run, and why the copy button exists.