Skip to content
AbilityBench

Wordless against worded

Non-verbal reasoning test that shows you what the words were costing

Twelve items, untimed and free, six of them in shapes and six in sentences — and behind them sit twelve logical structures, each one written in both forms, so the two halves of your score are the same reasoning with and without language in the way. The page scores and times the halves separately and then shows both twins of all twelve, which is the only honest way to demonstrate the claim a non-verbal paper is sold on. It also tells you when your two halves are too close to call, which on six items a side is most of the time.

  • 100% free
  • No signup
  • 12 items, untimed
  • 6 wordless, 6 worded
  • Both twins shown after

12 items, no clock. Behind them sit 12 logical structures written twice over — once in a sentence and once in shapes — and you get each structure in exactly one of its two forms, 6 of each, drawn at random.

What this measures that a plain reasoning test does not

Not whether you can reason — the pages next door do that better with more items. What this one isolates is the cost of the language: the same ordering, the same rule, the same cycle, posed once with words in the way and once without them. Afterwards you see both twins of all 12 structures side by side, which is the part that makes the point.

A warning about the arithmetic before you start: 6 items a side is far too few to tell two abilities apart in one person, and the page will say so rather than draw you a profile. What 6 items a side can show is the direction and the timing, and those are what it reports.

Non-verbal does not mean instruction-free

Each wordless item still carries one line telling you what to look for, exactly as an eleven-plus paper does. The reasoning material has no language in it; the rubric has to. That gap is the first limit on any claim that a figural test is culture-free, and it is easier to notice here than to argue about in the abstract.

Four options an item, answered with the number keys or by clicking. Shapes are told apart by outline, fill density, dot count and size, never by color, so nothing here depends on distinguishing one hue from another.

How the contrast is run

One structure, two forms, and only one of them per visitor.

  1. Take twelve items without knowing which twin you got

    Each item is either a sentence with four written options or a set of shapes with four shape options. What you are not told during the run is that the item you just answered has an identical twin in the other form — because seeing both would turn the second one into a memory test rather than a reasoning one. The assignment of six structures to each form is drawn from the run's seed, so a second run redraws it and shows you the other side of six structures.

  2. Read the two halves apart, and the interval on each

    The result gives wordless and worded as separate scores with a 95% Wilson interval on each, plus the median time an item on both. Six items a side is very little: only six right against none right produces intervals that fail to overlap, so the page computes that check on your own numbers and says plainly when the halves cannot be told apart. The timing is the softer figure and usually the more informative one, because the gap between the two medians is the reading.

  3. Open the side-by-side and see what each form demanded

    The review lists all twelve structures with both twins printed together, the item you were given highlighted, and one line on what each form asks that the other does not. That line is the content: the letter series that needs the alphabet backwards, the b-against-d reflection that literate adults have drilled for years, the grid that holds a constraint on screen while its worded twin makes you hold it yourself.

Technical specifications

Items12 scored — six wordless and six worded — drawn from 12 logical structures, each of which exists in both forms. No structure is shown to a visitor twice in one run
The twelve structuresTransitive ordering, denying the consequent, a two-condition conjunction, odd one out, two counters moving oppositely, a quarter-turn analogy, a reflection analogy, a three-step cycle, distribution of three, a two-rule series, a constant-ratio series and an analogy by degree
OptionsFour per item in both forms, so blind guessing averages 3.0 of 12. Wordless options are shapes rather than letters, because an option list in words would put the language back in the answer
Resolution of the contrastSix items a side. A 95% Wilson interval on six spans 61–100% for six correct and 0–39% for none, and no other pair of scores produces bands that clear each other — five right against one right still overlaps, so anything short of a clean sweep against a clean miss is a direction rather than a difference
How the shapes differOutline, number of sides, fill density on a three-step open–hatched–solid scale, a dot count from 1 to 6, relative size within a panel, rotation and reflection. No item is decided by color, so nothing here depends on telling one hue from another
TimingNone imposed. Response times are recorded and reported as a median per form, and the difference between those two medians is what the reading costs you — the one figure on the page that a wordless paper is actually claiming to remove
Reference figure≈ 3 IQ points per decade of rising raw performance when norms are held fixed — Flynn (1987), Massive IQ gains in 14 nations: What IQ tests really measure, Psychological Bulletin. It is quoted here because the rise is largest on exactly the figural items a non-verbal paper is made of, which is the strongest evidence against calling them culture-free
Reference figure for this testNone. The published figural batteries keep their norm tables under license, and a leaked table would be worse than none: with performance on these items rising generation on generation, an undated norm scores you against a population that has stopped existing

Frequently asked questions

Is a non-verbal reasoning test culture-free?

No, and the clearest evidence is a number rather than an argument. Raw performance on figural items rose by roughly three IQ points a decade across industrialized countries through the twentieth century, and the rise was largest on exactly these item types — matrices, series, analogies of shape. Something that changes that fast within one species and one gene pool is responding to schooling, print, diagrams and test familiarity, none of which is distributed evenly. What a wordless paper does remove is vocabulary, syntax and reading rate, which is a real and useful subtraction. Culture-free is a much bigger claim and it has never been supported.

Why do the wordless items still carry a line of English?

Because a figural item needs a rubric and no figural test has ever escaped that, the eleven-plus papers included. The reasoning material has no language in it — the shapes carry the ordering, the rule, the cycle — and the one line above them tells you what kind of answer is wanted. That line is the first and smallest limit on the comparability claim, and this page prints it rather than hiding it: a candidate who cannot read the rubric is at exactly the disadvantage the wordless items were supposed to fix.

Why only twelve items? That seems too few to measure anything.

It is too few, and the page says so in the results rather than pretending otherwise. Twelve is what the design costs: each structure has to be written twice and each visitor may see only one twin, so twelve items means twenty-four authored ones and six per half. This page is a demonstration of a difference, not a measurement of an ability — for an actual figural score with enough items behind it, the abstract-reasoning and matrix pages run twenty and upward, and the honest division of labor is that they measure and this one explains.

What are non-verbal reasoning tests actually used for?

Three settings, and they want different things from the same items. Selective school entrance at eleven uses them alongside verbal papers, on the reasoning that a child's reading age should not decide the whole result. Multinational hiring uses them because a figural item can be sat by candidates whose first languages differ without translating an item bank, which is a genuine practical advantage over a verbal test with the same aim. And assessment of children too young to read fluently, or of adults working in a second language, uses them because a verbal score there is partly a measure of the language. In all three the claim is comparability, not fairness in some larger sense.

I scored better on the worded half. Am I not a visual thinker?

The result does not support that reading, and neither does the underlying idea. Six items a side cannot separate two abilities in one person unless the split is six against nought, which the page checks and reports. Beyond the arithmetic, the visual-thinker framing does not survive contact with these items: two of the worded twins here are pure recall for a literate adult, since b against d and the compass points are reflections and rotations you drilled as a child. Doing better on those is your schooling showing up, which is the opposite of a fact about your reasoning style.

Are these the same as Raven's matrices?

No, and deliberately not. A matrix item is a specific figural form — a grid with one cell missing and rules running along rows and columns — and it is the form the copyrighted batteries are built from. Only one of the twelve structures here is a distribution grid, and it is here because its worded twin is a scheduling puzzle rather than because the grid needed rebuilding. The items on this page exist to be paired with sentences; if you want the matrix form done properly and at length, that belongs on the pages built for it.

Can I use this to prepare a child for the eleven-plus?

Not as practice, and possibly as an explanation. A selective-entrance paper is dozens of items under a tight clock in a specific publisher's house style, and twelve untimed items in nobody's style will not build the speed that paper is testing. What this page can do is show a parent what the non-verbal section is for and why it sits next to a verbal one — and the side-by-side makes that visible in about a minute, which no amount of prose about culture-fair testing does.

What taking the language out actually removes

Put the same logical structure in a sentence and in a set of shapes and the difference is narrower than the marketing suggests. The sentence adds vocabulary — knowing that wet is more than damp rather than different from it — and syntax, so that every and not have to be parsed before anything can be inferred, and reading rate, which under a clock is a score in its own right. Take those three away and what remains is not a purer reasoning: it is reasoning over a different notation. A dot count still has to be counted. A fill scale still has to be read as ordered rather than as three unrelated textures. A quarter turn has to be told from a reflection, which is a distinction that looks obvious once somebody has drawn it for you and is not obvious before. Two of the pairs here make the point in the other direction: the alphabet series and the compass analogy have worded twins that are pure recall for a literate adult, so the wordless versions are the harder ones, and the constant- ratio series is harder without words because a size has to be judged by eye instead of read off a numeral.

The three real uses of a figural paper all rest on comparability rather than on fairness. Selective school entrance at eleven pairs a non-verbal section with a verbal one so that a child’s reading age does not decide the whole result. Multinational hiring buys figural batteries because one item set can be sat by candidates whose first languages differ without translating and re-validating an item bank, which is a real saving and a real improvement over a translated verbal test. Assessment of children who cannot yet read fluently uses them for the same reason. None of that requires the items to be free of culture, and the strongest evidence that they are not is a figure this site is allowed to print: ≈ 3 IQ points per decade of rising raw performance with norms held fixed — Flynn (1987), Massive IQ gains in 14 nations: What IQ tests really measure, Psychological Bulletin — with the rise largest on precisely the abstract figural items that online tests copy. A score that moves that far in fifty years is not measuring something outside experience.

Two habits spoil most pages carrying this keyword. The first is the phrase non-verbal IQ, which converts a raw count on a handful of shape items into a number people repeat about themselves for years. No count produced here becomes an IQ point, and the box below says why. The second is rebuilding the matrix item for the fifth time in one site. A grid with a missing cell is a good item and it is already the subject of the progressive-matrix page and of the timed figural set, where it belongs and where there are twenty of them; and the pattern items split the same skill by whether the regularity is in number, shape or position. What has no other home is the comparison — so that is what this page is. Its natural neighbors are the worded halves of the same graduate sift: the true / false / cannot-say format, the five sub-scale profile and, for diagram items with physics in them, the trade-hiring format that is figural for a completely different reason.

Why there is no norm table on this page: Raven's Progressive Matrices norms

The item set and the norms are copyrighted and licensed. There is a second problem on top of the first: matrix reasoning is where the Flynn effect is largest, so even a norm table that had escaped into public circulation would be scoring today's visitors against a population that no longer exists.

A reaction time here is the interval between the frame that painted the stimulus and the timestamp the browser attached to your key, both read from the same monotonic clock. What neither can see is the display pipeline behind it, so on a 60 Hz screen roughly 16 ms of every figure below is the machine rather than you. That is the timing floor: two numbers closer together than that are the same number, and this page reports no precision it cannot support.

The figure it matters to here is the gap between the two medians, and that gap survives the floor by subtraction: whatever the hardware adds, it is added to both halves and disappears when one is taken from the other. What does not disappear is the sample: six items a side means a median built from six numbers, and a gap under a second and a half is reported as no gap at all for that reason.

This is a measurement exercise, not a clinical assessment. It reports what you did on this page against a stated reference and nothing more — it cannot establish a language difficulty, a visual-processing difficulty, or a thinking style. Only a qualified professional, working with more than a browser, can make that judgment.

Where the two halves are compared

Every number on this page is worked out by JavaScript running in the tab you are reading it in. Your answers, your reaction times and your score are never uploaded, logged or kept — which is also why the test carries on working after you disconnect from the network, and why nothing here can be held back behind an email address.

Both halves, their Wilson intervals and the timing gap between them are worked out from twelve answers held in this tab. The shapes are drawn as inline SVG by the page itself, so no image is fetched and there is nothing in a server log to say which twin you were shown — and because the assignment comes from a seed rather than from a record, a second run redraws it without either run knowing about the other.

Nothing is stored between runs, which is the whole reason the two halves can be lifted out as text: the summary you can put on the clipboard names the split, both intervals and the median gap, and it names no structure and no answer, so comparing this run with the next one is something you do rather than something the page does for you.