Overriding the first answer
Cognitive reflection test scored on how often the first answer won
Seven questions, free and untimed, each with a wrong answer that arrives before you have decided to think. You type your answer into an empty box, because a list of options would show you the answer you did not reach on your own. The number at the top is not how many you got right but how many of your answers were the lure itself — the particular wrong answer the question was written to produce — and a wrong answer that is not the lure is counted apart from one that is, because it means something different about how you got there.
- 100% free
- No signup
- 7 typed answers
- Lure rate, not just a total
- Score with the known ones removed
7 questions, each with a wrong answer that arrives on its own. You type your answer into a box rather than picking from a list, which is how these questions were written and the reason they work: a list would show you the answer you did not think of.
What the score is, and what it is not
Every question here has a declared lure — the specific wrong answer people give — as well as a solution. The headline figure is how often you handed over the lure, because that is the thing being measured: not whether you can solve the question, but whether the answer that showed up first went out unchecked.
A wrong answer that is not the lure is counted apart from one that is. It usually means the question was worked through and something slipped, which is a different event from not working it through.
The first 3 you may already know
The three published items open the run and they have been reprinted for twenty years, so some visitors arrive already holding the answers. There is nothing to be done about that except measure it: after the score you will be asked, item by item, which ones you had met before, and the report then gives the score again over only the ones that were new to you. The question comes afterwards on purpose — asked first, it would send you hunting through memory instead of through the problem.
No clock on any of them and no way to go back to one. Each item names the unit it wants, so “5”, “5 cents” and “$0.05” all count the same.
How to read a lure rate
One rate, two subscores that must be read apart, and a question you are asked only after the number exists.
Type an answer rather than choosing one
The box is empty and it stays empty until you fill it. That is not a styling decision: on a multiple-choice version of these questions the correct answer sits on screen next to the tempting one, and a visitor who never generated it can still recognize it, which destroys the thing being measured. Each question names the unit it wants, so five, five cents and $0.05 all read the same, and an answer the parser cannot get a number out of is recorded as unreadable rather than as an error. An empty box is a skip, and a skip is not counted as a lure.
Read the rate before the total
The headline divides your lure answers by the questions you answered. Two people who each got three of seven right are in completely different positions if one of them gave the lure four times and the other was wrong four different ways: the first went with the answer that came to mind, the second worked and slipped. That distinction is the construct, and it is why every question here declares its lure in advance instead of only its solution. The total is printed too, one line down, because that is the number other sites stop at.
Then mark which questions you had met before, and read the second score
The three published questions have been reprinted for twenty years, so some visitors arrive already holding the answers, and recall is not reflection. After the result you are asked item by item whether each was new to you, and a second panel then gives the score over only the new ones. The question comes after the number rather than before it on purpose — asked first, it would send you searching your memory instead of the problem — and nothing is held back while you answer it, since the full result is already on screen.
Technical specifications
| Questions | 7 in a fixed order: the three published bat-and-ball, machines and lily-pad items first, then 4 written for this page — a rope ladder on a rising tide, four eggs in one pan, a snail on a wall, and two successive 20% reductions |
|---|---|
| Answer entry | A free-text box parsed for a number, so “5”, “5 cents” and “$0.05” all count as five, and “none” counts as zero. Anything with no number in it is recorded as an unreadable answer; an empty box is a skip and is reported separately from a wrong answer |
| Outcomes per item | Three, not two: correct, the declared lure, or wrong by some other route. Every question carries its lure in the item data — 10 cents, 100 minutes, day 24, 3 rungs, 32 minutes, day 12, 40% — so the classification is exact rather than inferred |
| Headline figure | Lure answers divided by questions answered, as a percentage. Below it: the count correct, and the split of every wrong item into lure, other, and skipped |
| Exposure question | Asked after the score, one control per item, and never before it. The report then repeats the count over the items marked new to you and over the items you had met, so the effect of prior exposure on your own run is visible rather than assumed |
| The one published figure | 1.24 of 3 was the mean on the three original items among students at US universities, several thousand respondents across university samples — Frederick (2005), Cognitive Reflection and Decision Making, Journal of Economic Perspectives. It is printed beside your score on those three only, never beside the seven-item total, and it supports a comparison rather than a percentile because no individual-level spread was published with it |
| Why nothing is placed against chance | There is no guessing floor, and that is the point of a typed answer. An empty box has no guessing floor, so unlike a multiple-choice run there is no probability of reaching a score by luck to quote — which is why this test has always been asked open-ended |
| What is timed | Each question is clocked from its first painted frame until you submit, and you may sit on one as long as you like. The report prints the median on correct answers against the median on lure answers, with a 95% Welch interval on the difference, because overriding an answer costs time and that cost is measurable within one visitor |
Frequently asked questions
Why does the ball cost 5 cents and not 10?
Because at 10 cents the pair would come to $1.20, not $1.10. The bat has to cost a dollar more than the ball, so if the ball is 10 cents the bat is $1.10 and the total is $1.20. Put the ball at 5 cents and the bat is $1.05: the difference is exactly a dollar and the two add to $1.10. The reason 10 arrives first is that $1.10 splits into a dollar and ten cents without any arithmetic at all, and that split is precisely the one that throws away the constraint the question was built around.
I already knew all three of the famous questions. Is my run worthless?
No, but the total from those three is a memory score and the page treats it as one. That is why four more questions are here, why you are asked which items were familiar, and why a second panel reports the count over the new ones only. If you knew the first three and were new to the last four, the honest figure is the one from those four, and it is printed on its own regardless of how you mark anything. What no version of this test can do is un-see an item: prior exposure is a permanent property of a visitor and the only defensible response is to measure it rather than to hope.
Is a high score here the same as being intelligent?
No — it is a narrower thing, and the narrowness is what makes it interesting. These items are solvable by anybody who does the arithmetic; what they test is whether the arithmetic gets done at all once an answer is already present. That is a different question from how much you can hold in mind or how quickly you spot a pattern, and it is measured with seven questions, which is far too few to say much about a person. Numeracy also gets in the way of the original three, since all three need a small calculation, and that is one of the reasons an extension was proposed at all.
Why are there four extra questions instead of just the three?
Because three items give a score with four possible values, and two of those values are almost everybody. A three-item scale cannot separate people, cannot support any of the reliability arithmetic you would want, and is unusable the moment a visitor recognizes even one item. The published argument for extending it is that the construct is not tied to the original arithmetic — the same override can be provoked by questions that need almost no calculation, which is what our rope-ladder item does. Those four are ours rather than any published set, so no published mean applies to the seven-item total, and the page says so where the total is printed.
Does taking longer improve the score?
Usually yes, and the page measures that rather than treating it as cheating. Overriding an answer that has already arrived takes time, so correct answers here tend to have longer durations than lure answers, and the report prints both medians with a confidence interval on the difference. There is no time limit and no penalty for a slow answer, because the construct is whether you check rather than how fast you are. What a long duration cannot do is rescue a lure: an item answered with the tempting number after ninety seconds counts exactly as an item answered with it after three.
Why is a wrong answer that is not the lure counted separately?
Because it almost always means the opposite thing. Giving 12 on the snail item is the lure — one meter a day up a twelve-meter wall — and it means the problem was not opened. Giving 10 means it was opened, the daily net was worked out, and the final climb was handled like all the others; that is an arithmetic slip inside a reflective attempt. Folding those two together produces a total that cannot distinguish someone who never engaged from someone who engaged and miscounted, which is exactly the distinction the test exists for.
Is this the same test as the one in the original paper?
The first three questions are the published ones and the other four are not. That means your score on the three is directly comparable to the figure printed beside it, and your score on all seven is comparable to nothing outside this page. Two smaller departures are worth knowing about: the original was administered on paper inside a longer questionnaire, and each question here names the unit it wants — cents, minutes, a day number — so that a right answer written in the wrong unit is not scored as wrong. Neither changes what is being asked, and both are stated here rather than left for you to discover.
What the three questions measure, and what twenty years of reprinting did to them
The idea behind these items is narrower than the phrase “cognitive reflection” suggests, and the narrowness is the point. Each question is built so that a wrong answer presents itself immediately, is easy to check, and is wrong for a reason anybody can see once they look. So the item does not test whether you can solve it — nearly everyone can — but whether the checking happens when an answer is already sitting there. That was the case made in Frederick (2005), Cognitive Reflection and Decision Making, Journal of Economic Perspectives, which reported a mean of 1.24 of 3 across students at US universities, several thousand respondents across university samples. That number is a landmark rather than a rank, and this page prints it only beside your score on those same three items: it is a student mean, not a population one, so being above it does not place you above the general public, and with no individual-level spread published alongside it there is no percentile to compute from it at all.
The problem the intervening twenty years created is exposure. These three questions are among the most reprinted on the internet, and a visitor who already knows the ball costs five cents is not overriding anything — they are recalling. Almost every version of this test online ignores that and reports a total anyway, which is the single biggest defect in the genre: the score of a well-read visitor and the score of a careful one are indistinguishable. This page handles it in two ways. Four additional questions are asked that are ours rather than anybody's published set, written on the argument that the construct survives outside the original arithmetic items — Thomson & Oppenheimer (2016), Investigating an alternate form of the cognitive reflection test, Judgment and Decision Making is where that argument was made and tested, and our four are not its four. And after the result is on screen, you are asked which items you had met before, with the count then repeated over the new ones. Asking that first would have been worse than not asking: a visitor primed to consider familiarity searches memory instead of the problem.
The second departure from how these items are usually presented is the scoring itself. Three outcomes are recorded per question rather than two, because the specific wrong answer and any other wrong answer mean different things — one says the question was never opened, the other says it was opened and something slipped. That is also why the answer box is empty rather than a list: seeing the correct answer beside the tempting one lets recognition stand in for reflection, and the resulting score measures neither. The nearest page to this one is the logical reasoning test, where the tempting answer is a conclusion that happens to be true, and the deductive reasoning test puts the same conflict inside a single rule. If it is the arithmetic rather than the override that interests you, the numerical reasoning test allows a calculator and measures whether you can find the right two numbers; the pattern recognition test asks the same question about sequences, and the mechanical aptitude test is where a first intuition about levers and gears gets checked against how they actually behave.
A reaction time here is the interval between the frame that painted the stimulus and the timestamp the browser attached to your key, both read from the same monotonic clock. What neither can see is the display pipeline behind it, so on a 60 Hz screen roughly 16 ms of every figure below is the machine rather than you. That is the timing floor: two numbers closer together than that are the same number, and this page reports no precision it cannot support.
Here the clock is doing real work rather than sitting in a footnote, because the difference between the time your correct answers took and the time your lure answers took is one of the two figures this page reports. Those durations run to tens of seconds, so the floor is thousands of times smaller than the effect, and both are printed to a tenth of a second so that no precision is claimed which the browser cannot deliver.
This is a measurement exercise, not a clinical assessment. It reports what you did on this page against a stated reference and nothing more — it cannot establish impulsivity, a thinking style, or how carefully you decide anything away from these seven questions. Only a qualified professional, working with more than a browser, can make that judgment.
Where the seven typed answers go
Every number on this page is worked out by JavaScript running in the tab you are reading it in. Your answers, your reaction times and your score are never uploaded, logged or kept — which is also why the test carries on working after you disconnect from the network, and why nothing here can be held back behind an email address.
The strings you type are held in this tab so the review can show you what you wrote next to what the answer was, and they end when the tab does. Nothing is uploaded, nothing is recorded about which items you marked as familiar, and no analytics event carries an answer or a duration — which is also why a second visit cannot tell you how the first one went unless you copied the summary yourself.