Skip to content
AbilityBench

Validity, not truth

Logical reasoning test scored on whether it follows, not whether it is true

Sixteen short arguments with their conclusions already written out, free, with no clock on any of them and nothing held back at the end. Eight of them follow from their premises and eight do not; independently of that, eight carry a conclusion that is true in the world and eight carry one that is false. That crossing is the whole design: it lets the page report how much your accuracy moved when agreeing with the conclusion stopped being a safe guide, and it puts the four arguments that are true and still invalid in a panel of their own.

  • 100% free
  • No signup
  • 16 arguments, 2 answers each
  • 4 true but not entailed
  • No time limit

16 short arguments, each with its conclusion already written out for you. One question on every one of them: does the conclusion follow from the two premises above it? Not whether the conclusion is correct, and not whether the premises are — whether the step between them holds.

The instruction that decides most of your answers

Treat both premises as true for the length of that item, however wrong they sound. Some of them are wrong on purpose: an argument can carry you correctly from a false starting point to a false ending, and that argument is valid. Equally, a conclusion you know to be true can sit under premises that never delivered it, and that argument is not.

8 of the 16 let you answer from either instinct and get it right. On the other 8 the two answers differ, and 4 of those are the hardest version: a conclusion that is true, under premises that never produced it.

2 worked items come first, with the reasoning shown after each one. They use two steps that appear nowhere in the 16 scored items, so nothing you learn from them hands you a scored answer.

Answer by clicking, or with 1 for follows and 2 for does not follow. No item has a time limit; the seconds are recorded but nothing is cut off.

How to read the belief-bias figure

One headline number, one panel of four arguments, and the table that shows where your answers actually went.

  1. Answer as though both premises were handed to you as facts

    Several arguments start from something false — that glass floats, that airliners cannot leave the ground — and that is not a mistake on the page. An argument can carry you correctly from a false premise to a false conclusion, and it is still valid, because validity is a property of the step and not of the starting material. Judging those items on whether the conclusion is true is exactly the error being measured, so the instruction to grant the premises is the instruction that decides most of your score.

  2. Read the headline as a change in accuracy, not as a score

    The figure at the top is your accuracy on the eight arguments where logic and belief pointed the same way, minus your accuracy on the eight where they collided. Both halves run the same four steps in the same numbers, so difficulty is held constant and what is left is the pull of the conclusion. With four items in each cell the intervals are wide, and the panel says so: read a gap under about thirty percentage points as unproven rather than as absent, and take a second run on a different day rather than trusting a small one.

  3. Then look at the four arguments that were true anyway

    The second panel lists the four whose conclusions you would accept without argument — that no cat is a dog, that some surgeons work nights — and none of which their premises delivered. Each one shows what you answered and the counterexample that breaks it: a substitution that leaves the premises true and turns the conclusion false. Refusing four of four there is the single most informative outcome on the page, and refusing none of four is the ordinary one.

Technical specifications

Arguments and design16 scored arguments in a crossed set: valid or invalid, crossed with a conclusion that is true or false in the world, four to a cell. Order is reshuffled from a seed on every run; the items themselves are fixed
Forms matched across the splitThe two invalid cells run the same four steps — a shared middle term that never gets shared, twice; two premises that both begin with “some”; and an “or” taken to exclude its other branch. The two valid cells are matched the same way, so a difference between cells is belief and not difficulty
Answer formatTwo buttons, follows or does not follow, plus keys 1 and 2. A coin scores 8 of 16, so the result panel prints the exact one-sided binomial probability of your count arriving by guessing, and a Wilson interval on your accuracy
Worked arguments first2 worked arguments with the reasoning shown after each, built on a universal-plus-denial step and on a term that widens between premise and conclusion. Neither step appears in the 16, so nothing learned from them is an answer you would otherwise have had to find
Feedback during the runNone after the warm-up. A verdict per argument would teach the crossing by about item six, and the two halves would stop being comparable — the measurement would be gone rather than merely noisier
What is timedThe clock on an argument starts once its text has finished painting and stops on the click or keypress that answers it; nothing expires while you think. The medians on the two halves are printed as context: extra seconds on a conflict item are where a first impression got checked instead of acted on
Reference figureNone printed, and the reason is not modesty. There is no population norm for 'a verbal reasoning test' or 'a pattern recognition test', because those name a question format rather than an instrument. The score depends entirely on how hard the page made its own items, so any average quoted would describe this page's difficulty setting and nothing about people. The commercial tests in these formats do have norms, and those are restricted — see the hiring-assessment entry.
Paradigm sourceEvans, Barston & Pollard (1983), On the conflict between logic and belief in syllogistic reasoning, Memory & Cognition — cited for the crossed design and for nothing else. Its acceptance rates were collected on paper with a different item set, so none of them is printed here and your result is not scored against them

Frequently asked questions

What is the difference between a valid argument and a true conclusion?

Validity is a property of the step from premises to conclusion; truth is a property of a statement. An argument is valid when there is no way for its premises to be true and its conclusion false at the same time, which says nothing about whether the premises are in fact true. So “everything made of glass floats; window panes are glass; therefore window panes float” is valid and ends in something false, while “all cats are mammals; all dogs are mammals; therefore no cat is a dog” is invalid and ends in something true. Formal reasoning tests score the first property, and almost everyone who loses points is answering the second.

Why does the test include premises that are obviously false?

Because a set of arguments with true premises cannot tell the two skills apart. If every conclusion that follows is also true and every conclusion that does not is also false, then answering from what you know about the world scores exactly as well as reasoning, and the test measures nothing it was built to measure. False premises are the only way to produce an argument that is valid and ends somewhere false, which is one of the four cells here. The eight items in the two conflict cells are what give the page a number that a guess from general knowledge cannot reach.

Why is “no cat is a dog” marked as not following?

Because the premises put cats inside the mammals and dogs inside the mammals, and nothing else. Two classes that both sit inside a third can overlap completely, partly, or not at all, and the premises do not say which. The test of a step is whether the same shape ever fails: replace “dogs” with “retrievers” and you get “all dogs are mammals, all retrievers are mammals, therefore no dog is a retriever”, which is false with both premises still true. Cats being distinct from dogs is a fact you brought with you, and this item is scored on where it came from.

Does “some” mean “not all” on this test?

No — “some” means at least one, and it leaves open that all of them do. In ordinary speech saying “some of the specials are vegan” hints that the rest are not, which is a conversational implication rather than part of what the sentence claims. The scored items never depend on the difference: no item is valid only if “some” excludes “all”, and none is invalid only because it does. If you read “some students are vegetarians” as “not all students are vegetarians”, every answer on this page still comes out the same.

Does “or” mean one or the other but not both?

Not here, and that is one of the four steps being tested. In logic “A or B” is satisfied by A, by B, and by both together, so learning that one branch happened tells you nothing about the other. Two of the sixteen arguments turn on this: “the parcel went by courier or by post, it went by courier, so it did not go by post” is invalid even though it sounds airtight, because the premise never promised exclusivity. English does have an exclusive “or”, usually signaled by “either … or … but not both”, and no item on this page uses that wording.

Why only two answers instead of a “cannot say” option?

Because “does not follow” already is the cannot-say answer, and adding a third option would split it in two. On a passage-based verbal test the three responses answer a different question — whether the passage asserts the claim, denies it, or is silent — and “cannot say” is the silent case. Here you are given the premises in full, so silence is not available: either they force the conclusion or they leave it open, and leaving it open is precisely what the second button says. The cost of two answers is a 50% chance level, which the result panel handles by printing the probability of your total arriving by luck.

Is a low score on the conflict half a sign of poor logic?

It is a sign that on these eight arguments your knowledge of the world was doing the work, which is a description of eight arguments rather than of you. The effect it measures is close to universal and gets larger, not smaller, when people are told to ignore what they know — that is why the design exists rather than a simple instruction. Four items in each cell also leave a wide interval, so a gap of one or two items is noise. What the panel is good for is the specific list: seeing which arguments you accepted, and reading the substitution that breaks each of them, is the part that transfers.

Belief bias, and why a reasoning test has to work against what you already know

Formal logic separates two things that ordinary reading fuses. Whether a conclusion is true is a question about the world; whether it follows is a question about the argument, and the second can be settled without knowing anything about the subject matter. Every assessment built on syllogisms scores the second, and the reason they keep working is that people cannot easily stop answering the first. The finding has a name and a fixed shape: acceptance of a conclusion goes up when the conclusion is believable, whether or not the argument supports it, and the effect is largest exactly where the two come apart. The study that isolated it by crossing validity against believability inside one item set is Evans, Barston & Pollard (1983), On the conflict between logic and belief in syllogistic reasoning, Memory & Cognition, and the sixteen arguments here are built on the same crossing rather than on its items.

What most online reasoning tests do instead is ask arguments whose conclusions are all plausible, which quietly removes the measurement. If every valid argument also ends somewhere true, a visitor who ignores the premises entirely and rates conclusions on general knowledge scores the same as a visitor who reasons — and a score that two completely different procedures produce equally well is not measuring either. The second common failure is subtler: including conflict items but reporting one total, so a person who was perfect on the easy half and lost every conflict item shares a score with someone who was mediocre throughout. Splitting the two halves costs nothing and is the difference between a number and a diagnosis of where the number came from. If you want the same distinction narrowed to one rule applied to one case, with the invalid steps named as you make them, that is the deductive reasoning test; the five-part appraisal that grades inference, assumption, deduction, interpretation and argument evaluation separately is the critical thinking test; and where the wrong answer arrives quickly and feels finished rather than following from any premise at all, that is the cognitive reflection test.

No population figure is printed beside your score, and the reason is worth stating plainly. There is no population norm for 'a verbal reasoning test' or 'a pattern recognition test', because those name a question format rather than an instrument. The score depends entirely on how hard the page made its own items, so any average quoted would describe this page's difficulty setting and nothing about people. The commercial tests in these formats do have norms, and those are restricted — see the hiring-assessment entry. The commercial tests written in this format do carry norms, and those norms belong to their publishers and were collected under a stopwatch with an item bank nobody else may use. So what this page reports is a count, an interval around it, the probability of that count arriving by chance, and a within-visitor difference between two matched halves — all four of which are meaningful without any comparison group at all. For the same judgment applied to short passages, where the third response option means the passage is silent rather than that the claim is false, the verbal reasoning test is the closer format, and for reasoning with no sentences in it at all, the matrices page removes language from the path entirely.

A reaction time here is the interval between the frame that painted the stimulus and the timestamp the browser attached to your key, both read from the same monotonic clock. What neither can see is the display pipeline behind it, so on a 60 Hz screen roughly 16 ms of every figure below is the machine rather than you. That is the timing floor: two numbers closer together than that are the same number, and this page reports no precision it cannot support.

Nothing on this page is scored on speed, so that floor is nowhere near large enough to matter to it: the durations printed beside the two halves run to seconds, and they are shown to a tenth of a second rather than to the millisecond for exactly that reason. It is stated because the clock is the same one every timed page here uses, and because a number is only as good as the account of where it came from.

This is a measurement exercise, not a clinical assessment. It reports what you did on this page against a stated reference and nothing more — it cannot establish a reasoning impairment, a learning difficulty, or how sound your judgment is outside these sixteen arguments. Only a qualified professional, working with more than a browser, can make that judgment.

Where the sixteen answers stay

Every number on this page is worked out by JavaScript running in the tab you are reading it in. Your answers, your reaction times and your score are never uploaded, logged or kept — which is also why the test carries on working after you disconnect from the network, and why nothing here can be held back behind an email address.

The four-cell table and the panel of true-but-invalid arguments are drawn from sixteen answers and sixteen durations held in an array in this tab, for as long as the result is on screen. Nothing is uploaded, no analytics event carries an answer, and a second visit has no way to know what the first one did — which is also why comparing two runs means keeping the copied summary yourself.