Skip to content
AbilityBench

One rule, one case

Deductive reasoning test that names the step you got wrong

Twenty items, free and untimed, each giving you a rule of the form “if this, then that”, one fact about a single case, and three answers to pick from. On ten of them the rule licenses a conclusion and on ten it settles nothing, and choosing anyway is not carelessness but one of two specific errors that have had names for a century. The report gives you four rows, one per position a conditional can be questioned from, and says which of the two errors is yours rather than adding them together. Two four-card puzzles follow the twenty.

  • 100% free
  • No signup
  • 20 items, 4 named forms
  • 2 card puzzles
  • 3 answers an item

20 items, built from five rules of the form “if this, then that”. Each item gives you the rule, then one fact about a single case, and asks what the rule lets you conclude about that case. Three answers each, and one of the three is always that the rule settles nothing here.

Where the 10 unanswerable items come from

A rule of that shape is a one-way license. It promises the consequence whenever the condition holds, and it promises nothing about the cases where the condition is absent, because it never said the condition was the only way to get there. Half of the items put you in exactly that position, and the honest answer is that nothing follows.

Choosing otherwise is not carelessness — it is one of two specific mistakes with names, and the report tells you which of the two is yours rather than adding them together.

After the 20 items there are two card puzzles, the same rule stated twice in different clothing. They are untimed, they are not part of the score above, and the second of the two is usually much easier than the first for a reason worth knowing.

Click an answer, or press 1, 2 or 3. Take as long as you want on any item — the seconds are recorded and nothing is cut short.

How to read the four-position report

Two headline counts that move independently, four rows, and two card puzzles that are the same problem twice.

  1. Use the rule and nothing else

    Every item names a rule and then one fact, and the answer has to come from those two lines alone. Whether porch lights usually come on with an alarm, or whether a heavy parcel would realistically go by courier, is knowledge you brought and it is not part of the item. Four of the five rules are ordinary enough that plausibility and entailment agree, which is deliberate: the fourth position on each rule is where they part, and that is the item worth being careful on.

  2. Read the two headline counts as a pair

    One count is how many of the ten unanswerable items you refused; the other is how many of the ten licensed conclusions you actually drew. They move independently on purpose, because either one alone can be gamed. Answering that nothing follows every single time gives you a perfect first count and a zero second one, and treating the rule as though it worked in both directions gives you the reverse. Only both together say anything, and the four rows below split them into which position each answer came from.

  3. Then compare the two card puzzles against each other, not against your score

    The two four-card problems after the twenty are formally the same claim: one about vowels and even numbers, one about beer and age, with the same four positions in the same order. Almost everybody finds the second easy and the first hard, and the interesting output is whether you did too. They are untimed, they sit outside the scored block, and turning a card you did not need counts against the answer, because the question is which cards must be turned rather than which might be interesting.

Technical specifications

Items and rules20 scored items: 5 rules of the form “if this, then that” — a courier threshold, a pass list, a porch light, a reference shelf and a safety check — each asked from all 4 positions. Rules, positions and options are reshuffled from a seed on every run
The four positionsCondition present, which licenses the consequence; consequence absent, which licenses the absence of the condition; consequence present, which licenses nothing; condition absent, which licenses nothing. Five items each, scored as four separate rows
Option orderShuffled per item rather than fixed. “Nothing follows” is the right answer on 10 of the 20, so pinning it last would let a visitor score half the run by always choosing the bottom option; the error type is recovered from what the chosen option asserted, not from where it sat
Guessing floorOne in three. Twenty blind guesses average 6.7 items correct; the report gives the chance of reaching your own total that way, computed from the binomial rather than approximated, and a Wilson interval beside it
Selection task2 versions of the four-card problem, abstract first and then the same rule as an age check, with 4 cards each and free multi-select. Untimed, outside the scored block, and marked as a whole selection: an unnecessary card turned makes the answer wrong
Published comparison≈ 10% of participants pick the right two cards on the abstract version — Wason (1968), Reasoning about a rule, Quarterly Journal of Experimental Psychology. It is written, self-paced and needs no examiner, which is why a browser version is the same measurement and why this is the one figure on the page with a population behind it; the same rule in a familiar social setting is solved by most people, per Griggs & Cox (1982), The elusive thematic-materials effect in Wason's selection task, British Journal of Psychology
Explained items first2 worked items on a kettle rule, one from each side of the valid split, with the tempting answer named before the scored run starts. The kettle rule appears in none of the 20
Reference figure for the twentyNone. There is no population norm for 'a verbal reasoning test' or 'a pattern recognition test', because those name a question format rather than an instrument. The score depends entirely on how hard the page made its own items, so any average quoted would describe this page's difficulty setting and nothing about people. The commercial tests in these formats do have norms, and those are restricted — see the hiring-assessment entry.

Frequently asked questions

The porch light is on, so is the alarm armed?

No — that is the error the page is named after, and the rule does not support it. “If the alarm is armed, the porch light is on” promises one route to the light being on and never claims it is the only one; the light might be on a timer, or somebody might have flipped the switch. Going from the consequence back to the condition is affirming the consequent, and it is the most common wrong answer on any conditional test. It feels safe because most real rules do work in both directions, and the ones that do say so.

Is modus tollens really valid? It feels like a trick.

It is valid, and the reason is that the rule rules out exactly one combination. “If the parcel is over 2 kg, it goes by courier” forbids a heavy parcel that did not go by courier, and nothing else. So being told a parcel did not go by courier settles that it was not over 2 kg, because the other possibility has been excluded by the rule itself. Around a third to a half of people decline this step on a written test, usually by answering that nothing follows — which is the rule read forwards only, as a one-way permission rather than as a constraint that also bites from the far end.

Why does “if” not mean “if and only if”?

Because the two make different promises, and only one of them is written on the page. “If A then B” forbids A-without-B; “A if and only if B” also forbids B-without-A, which is a second, stronger claim. Everyday speech leans on context to supply the stronger reading — a parent saying “if you finish your homework you can go out” usually means only then — and formal reasoning strips that context out. Every rule on this page is the weaker one, and half the items are positions where the difference between the two decides the answer.

Why do most people fail the letters-and-numbers cards and solve the beer one?

Because the second version turns a logic problem into a rule-enforcement problem, which people are much better at. The two puzzles are formally identical: the same conditional, the same four cards, one showing the condition, one showing its absence, one showing the consequence and one showing the absence of the consequence. In the abstract form roughly one person in ten picks the two that can reveal a violation. Dress the same rule as a drinking-age check and most people get it, because looking for a cheat is a task with obvious targets. That gap is the single most reproduced result in the reasoning literature.

How is this different from a logical reasoning test?

Deduction is one part of logic, and this page is the narrow one. Here every item is a general rule applied to a specific case, and the two invalid steps have names that the report uses. A broader logical reasoning set adds syllogisms, quantifiers and disjunctions, where the pull towards a wrong answer comes from the conclusion being true in the world rather than from the shape of the rule — that is a different measurement and it is on its own page. If you only want to know which of affirming the consequent and denying the antecedent is your habit, this is the page that answers it.

Do employers' deductive reasoning tests use this format?

The item type is the same and almost nothing else is. A publisher's deductive battery runs under a strict per-item time limit, draws from an item bank that is kept out of circulation precisely so it cannot be practiced, and reports a percentile against a norm group of applicants for comparable roles. This page has no time limit, twenty items you can read as slowly as you like, and no norm group at all. It is useful for meeting the four positions and finding out which of them you get wrong; it will not produce a number that converts into one of theirs.

Why is turning all four cards marked wrong?

Because the question is which cards you must turn, and turning everything answers a different one. Two of the four cannot show a violation whatever is on the back — the card that fails the condition and the card that already satisfies the consequence — so turning them buys nothing, and a selection that includes them has not identified what the rule forbids. The original scoring treats a superset as incorrect for exactly this reason, and this page keeps that rule so the result stays comparable to the published solution rate.

The two inferences a conditional grants, the two it refuses, and why the refusals are the test

A rule of the form “if A then B” makes exactly one prohibition: there is no A without B. Everything a deductive test can ask follows from that. Told A, you may write B down. Told not-B, you may write not-A down, because an A there would have needed a B and there is not one. Told B, you have learned nothing about A, since the rule never said B arrives only that way. Told not-A, you have learned nothing about B, for the same reason read from the other end. Two of the four positions license a conclusion and two do not, and it is the two refusals that discriminate: endorsement of the first is near-universal on written tests, while the other three vary enormously between people, which is why a single total across all four hides most of what happened.

The two wrong steps are old enough to have Latin names and specific enough to be worth recognizing in the wild. Reading the consequence back into the condition — the light is on, so the alarm is armed — is affirming the consequent, and it is the shape of most confident wrong diagnoses from a single symptom. Reading the absence of the condition into the absence of the consequence — the alarm is off, so the light is off — is denying the antecedent, and it is the shape of most confident wrong conclusions from a missing cause. Both feel fair, because most rules people meet really do run in both directions and the ones that do announce it. What almost no online test does is report which of the two you make: they are opposite errors, they respond to different corrections, and almost nobody makes them equally. The wider set that adds syllogisms and quantifiers, where the pull comes from the conclusion being true rather than from the rule's shape, is the logical reasoning test, and the five-sub-test appraisal used in law and consulting hiring is the Watson-Glaser format practice page.

The card puzzle after the twenty items is here because it is the one part of this subject with a number that survives the trip into a browser. It is written, self-paced and needs no examiner, so asking it on a web page is the same act as asking it on paper — ≈ 10% of participants select the two cards that can reveal a violation, per Wason (1968), Reasoning about a rule, Quarterly Journal of Experimental Psychology. What does not survive is the content: the identical rule presented as a check on who may drink is solved by most people, which is the result reported in Griggs & Cox (1982), The elusive thematic-materials effect in Wason's selection task, British Journal of Psychology. That pair is the strongest available argument that these tests measure something narrower than they appear to, since the same person can fail the abstract form and solve the social one in the same minute. It is also why the twenty items above use everyday rules rather than letters and numbers, and why nothing here is offered as a general verdict on anybody's reasoning. The nearest sibling that asks you to hold a rule against a passage rather than against a case is the reading comprehension test, and the page that catches the answer you gave without applying any rule at all is the cognitive reflection test.

A reaction time here is the interval between the frame that painted the stimulus and the timestamp the browser attached to your key, both read from the same monotonic clock. What neither can see is the display pipeline behind it, so on a 60 Hz screen roughly 16 ms of every figure below is the machine rather than you. That is the timing floor: two numbers closer together than that are the same number, and this page reports no precision it cannot support.

No score on this page depends on that floor: the per-position medians are in seconds and are printed to a tenth of a second, and the card puzzles are not timed at all because the published solution rate they are compared against was not. The clock is stated because it is the same clock, and because the seconds an item took are shown next to how you answered it.

This is a measurement exercise, not a clinical assessment. It reports what you did on this page against a stated reference and nothing more — it cannot establish a reasoning disorder, a learning disability, or how you will do on an employer's assessment next week. Only a qualified professional, working with more than a browser, can make that judgment.

Where the twenty answers and the card picks live

Every number on this page is worked out by JavaScript running in the tab you are reading it in. Your answers, your reaction times and your score are never uploaded, logged or kept — which is also why the test carries on working after you disconnect from the network, and why nothing here can be held back behind an email address.

The four-position table is built from twenty answers and twenty durations held in an array in this tab, and the card selections from two lists of card ids beside them. None of it is uploaded and none of it survives the tab closing, which is also why the card puzzles cannot be re-taken honestly on a second run: the answers are on screen, and the page has no memory of whether you had seen them.