Reasoning profile
Critical thinking test scored as five sub-scales, not one number
Twenty items with no clock on them, four each for five separable skills — reading how far a finding reaches, spotting the premise an argument never states, telling what follows from what is merely likely, reading a figure for exactly what it says, and judging whether a true statement is a reason for the claim at hand. Free, no account, every item explained afterwards. The output is a profile with a 95% interval drawn on each bar, and the page checks whether your strongest and weakest sub-scales genuinely separate before letting you read anything into the shape.
- 100% free
- No signup
- 20 items, untimed
- 5 sub-scales of 4
- Intervals on every bar
20 items, four options each, no clock. Five skills are measured 4 items apiece, and the result is the shape across the five rather than the total under them.
How far the evidence reaches
Given a finding, saying which claim it supports and which one it merely accompanies. The skill is stopping at the edge of the data instead of at the edge of plausibility.
The premise nobody stated
Finding the thing an argument needs to be true and never says. Not what would also be sensible to believe — what the argument collapses without.
What cannot fail
Separating conclusions that follow from the premises no matter what from conclusions that are merely very likely. The likely ones are where points go.
What the figures license
Reading a number for exactly what it says: a share is not a count, a relative fall is not an absolute one, and a sample is not a population.
Whether the reason bears
Judging whether a true statement is a reason for the claim in front of you. Most bad arguments are not false; they are about something else.
What four items per skill can and cannot tell you
A sub-scale of four is a proportion measured on four observations, and the interval around it is enormous: four right spans 51% to 100% and none right spans 0% to 49%. So this page draws the interval beside every bar and checks whether your strongest and weakest actually separate — on four items apiece, only a clean sweep against a clean miss does. The bars are worth reading. Small differences between them are not.
No clock, deliberately. A time limit would turn this into a reading-speed measure with a reasoning flavor, and several items are designed to be right only if you stop and do one division.
Answer with the number keys or by clicking. Options are reshuffled every run, so the position of a right answer carries no information. Every item is explained afterwards, whichever way you answered it.
How to read your own profile
Five skills, four items each, and one check on whether the shape means anything.
Answer twenty items with the five skills interleaved
Each item gives a short scenario and four options, one of which is right. The five skills rotate — one item from each, then the next round — because four deductions in a row teach their own trick and would inflate that sub-scale and nothing else. Options are reshuffled on every run, so the position of a right answer carries no information. Nothing is timed, and several items are only answerable if you stop and do one division.
Read the bars, then read the bands around them
The result draws five sub-scales with the score as a dark line and a 95% Wilson interval as a pale band. A band that straddles the 25% rule has not distinguished you from guessing on that skill. The bands are wide on purpose and they are wide because four items is four items: this is what a four-observation proportion honestly looks like, and every page that draws a five-spoke chart without them is drawing noise with confidence.
Check whether your strongest and weakest actually separate
Below the bars, the page compares the extremes of your own profile and says whether their intervals overlap. On four items apiece only a clean sweep against a clean miss clears the bar, so most runs come back flat — which is the true answer rather than a disappointing one. If you want the shape to mean something, run it three times on different orders and watch which skill comes last every time.
Technical specifications
| Items | 20 scored, four for each of the five skills, drawn in a fresh interleaved order every run. No practice block, because nothing is timed and there is no response mapping to learn |
|---|---|
| Options | Four per item, reshuffled each run, so blind guessing averages 5.0 of 20. One item in each skill is written so that the most plausible-sounding option is the wrong one |
| Timing | None. A limit would make this a reading-speed measure with a reasoning flavor; response times are recorded and shown per skill as information, and a slow correct answer scores exactly what a fast one does |
| Sub-scale resolution | A 95% Wilson interval on four items spans 51–100% for four correct and 0–49% for none correct. Those are the only two that fail to overlap, so a four-against-one profile is a flat profile — the check is computed on your own numbers rather than asserted here |
| Interruptions | Focus moving to another window is ignored; this tab being hidden with an item open sets that item aside and asks it again at the end. On an untimed reasoning item the page cannot tell thinking from looking it up, so it does not try to score one it did not watch |
| Reference figure, content dependence | ≈ 10% solve the abstract form of Wason's selection task — Wason (1968), Reasoning about a rule, Quarterly Journal of Experimental Psychology; most people solve the identical rule in a familiar social setting, per Griggs & Cox (1982). It is one of the few figures on this site collected under conditions a browser reproduces |
| Reference figure, overriding an intuition | 1.24 of 3 on the three-item Cognitive Reflection Test — Frederick (2005), Cognitive Reflection and Decision Making, Journal of Economic Perspectives. A student sample rather than a population, and those three items are among the most circulated questions online, so it is context and not a benchmark |
| Reference figure for this test | None exists. Critical thinking as a general ability has no public browser norm; the instruments that hold norm tables are hiring products whose figures are licensed and often cut by role, so the score is reported raw against a stated guessing floor |
Frequently asked questions
Why five scores instead of one critical thinking score?
Because a single score assumes the five move together, and the best-replicated result in this literature is that they do not travel across content, let alone across each other. The same logical rule that almost nobody handles in the abstract is handled by most people the moment it is dressed as a social situation — so a number claiming to capture your reasoning in general is claiming something the evidence does not support. Five sub-scales are a weaker claim and a truer one: they say what you did on four items of each kind, and the shape is a place to look rather than a trait.
My strongest and weakest sub-scales did not separate. Is the test broken?
No — that is the test working, and most runs come back that way. Four items give a proportion with a very wide interval on it, and two intervals that overlap are two scores you cannot tell apart at 95%. The page could hide that and print a five-spoke chart with a confident shape, which is what almost every free version does; it prints the bands instead. If the shape matters to you, the fix is repetition rather than arithmetic: three runs on different orders will show whether one skill really does sit last.
Why is there no timer?
Because the skills being measured are ones a clock actively damages. Four of the twenty items turn on stopping to work out a proportion — 90% of a 22% response rate, 78% of 62%, twelve defects in three hundred against five in a hundred — and a person who would have done that arithmetic under no pressure and skipped it under a clock has just been scored on their pressure tolerance. Timed reasoning is a real and separately useful measurement, and the version of these five under an exam clock is on the appraisal-practice page instead.
Some of these answers look debatable. Are they?
Two of the twenty could be argued and the rest could not, and the standard applied to all of them is stated in the explanation under each. The rule used throughout is that the right option is the one supported by the material given, not the one that is most likely true in the world — so an option can be correct about reality and wrong here. Where two options are both defensible, the item was rewritten rather than kept, because an item with two answers measures whether you share the author's taste. If you find one that still has two, the explanation will tell you which reading was used.
Can critical thinking be improved, or is this a fixed trait?
The parts this page measures respond to being named, which is why every item is explained afterwards whichever way you answered it. Reasoning that fails in the abstract routinely succeeds once the same structure appears in a familiar setting, and that gap is the strongest available evidence that what is missing is usually a way of seeing the problem rather than a capacity. The practical version of that: after a run, the useful move is not to retake it but to catch yourself doing the same thing on a real decision — accepting a rate without asking what the denominator was, or treating a nearby true fact as a reason.
How is this different from the hiring appraisal practice on this site?
Same five constructs, opposite jobs. The appraisal page reproduces an exam: its published section order, its five different response scales, a twelve-minute clock, and a report built around the five-point inference scale a candidate has to calibrate. This page throws all of that away and asks the same five questions as skills — four-option items, no clock, interleaved, and a profile with intervals. Take the exam version if you have an assessment day in the diary; take this one if you want to know which of the five is your weak one.
Is a critical thinking score the same thing as an IQ score?
No, and the difference is easiest to see in what each one survives. Abstract reasoning scores are famously stable within a person and famously unstable across generations; performance on these five collapses and recovers depending on how the same problem is dressed, which is not how a general capacity behaves. Nothing on this page converts to an IQ point, a percentile or a band, and no browser test does — the ones claiming otherwise are pricing a number they cannot produce.
Why critical thinking refuses to be one number
The oldest embarrassment in this field is also its most useful finding. Given a rule and four cards, and asked which to turn over to test it, ≈ 10% of participants pick the right two — Wason (1968), Reasoning about a rule, Quarterly Journal of Experimental Psychology. Dress the identical logical structure as checking whether drinkers are old enough and most people get it, which Griggs & Cox (1982) established and nobody has undone. Same rule, same required move, a forty-point swing in who can make it. That result is why this page reports a profile: if a single logical operation can be present and absent in the same person depending on the wrapping, a number claiming to summarize their reasoning in general is measuring the wrapping. Five sub-scales are a narrower claim, and narrower is what the evidence will carry.
The second thing worth knowing is how little four items resolve, because that is what everybody selling a critical-thinking profile hides. A score of four out of four carries a 95% interval running from 51% to 100%; a score of nought out of four runs from 0% to 49%. Those two barely fail to touch, which means a clean sweep against a clean miss is the only pair of sub-scales that separates on a run of this size — a four-against-one profile, which looks decisive on a radar chart, is statistically flat. The response is not to hide the sub-scales but to draw the bands and check the comparison, and to say plainly that a shape worth acting on comes from three runs and not from one. This page prints the check for your own numbers rather than a general reassurance.
What the five look like when they fail is concrete enough to recognize later. Reach fails when a study that chose its own groups gets read as though it had randomized them. Premise fails when a figure counting where orders arrive gets read as counting what caused them. Necessity fails when a conditional is run backwards from its consequence, which feels identical to the valid backwards move and is not. Figures fails when a rising percentage of a shrinking cohort gets read as a rising count. Bearing fails when a true statement about a different question gets counted as a reason. Related to the last one is the slower, more famous failure of overriding a first answer: the three-item reflection test averages 1.24 of 3 — Frederick (2005), Cognitive Reflection and Decision Making, Journal of Economic Perspectives, on students rather than a population, and the version here asks whether you had already met the items, because by now most people have. Beside it, the true / false / cannot-say format puts the same restraint on passages under a clock, the bare-premise items isolate the necessity sub-scale on its own, and the wordless version asks how much of any of it was language.
Wason selection task solution rate, abstract form
≈ 10% — share of participants who select the correct two cards on the abstract rule, among University students.
Wason (1968), Reasoning about a rule, Quarterly Journal of Experimental Psychology
The task is written and self-paced, so a browser version is the same measurement as the original — which makes this one of the few figures on this site that transfers cleanly. What does not transfer is content: dressing the identical rule up as checking drinkers' ages lifts the solution rate enormously. The abstract failure rate is therefore a fact about the abstract wording, not about the logic of whoever failed it.
A reaction time here is the interval between the frame that painted the stimulus and the timestamp the browser attached to your key, both read from the same monotonic clock. What neither can see is the display pipeline behind it, so on a 60 Hz screen roughly 16 ms of every figure below is the machine rather than you. That is the timing floor: two numbers closer together than that are the same number, and this page reports no precision it cannot support.
It is worth stating here only to be clear that it plays no part: nothing on this page is scored on time at all, the per-skill medians are printed in seconds as a curiosity, and the one figure a display could distort is not a figure anybody is being judged on.
This is a measurement exercise, not a clinical assessment. It reports what you did on this page against a stated reference and nothing more — it cannot establish a thinking style, a bias you are prone to, or a reliable habit of deciding well. Only a qualified professional, working with more than a browser, can make that judgment.
Where the twenty answers and the five intervals live
Every number on this page is worked out by JavaScript running in the tab you are reading it in. Your answers, your reaction times and your score are never uploaded, logged or kept — which is also why the test carries on working after you disconnect from the network, and why nothing here can be held back behind an email address.
The five sub-scales and their Wilson intervals are worked out from an array of twenty answers held in this tab, and the arithmetic runs on your machine rather than on ours. Running the test three times, which is what the profile actually needs, leaves nothing behind between runs either — the page has no memory of the previous one, which is why the copy button exists.