Skip to content
AbilityBench

Free 24-item test across three spatial sub-scales

Spatial reasoning test scored as three abilities, not one

Twenty-four items in an eight-minute budget, split evenly between turning a shape, folding a net into a cube and cutting through a solid, and reported as three separate counts with a confidence interval on each. Being quick at one of those three predicts surprisingly little about the other two, which is exactly why a single spatial score hides more than it says. Nothing is downloaded, nothing is stored, and every figure on screen is generated from numbers rather than scanned from a booklet.

  • 100% free
  • No signup
  • 24 items, 3 sub-scores
  • 8-minute budget
  • Figures drawn live

24 items in 8 minutes, in three kinds that arrive shuffled together: turning a flat shape, folding a net into a cube, and cutting a solid with a flat plane. One clock covers the whole set, so an item you dwell on is an item you may not reach.

Turning

A five-square shape and four candidates. One is it, turned; three are its mirror image, also turned.

Folding

Six squares in a strip or a cross, and the question of which face lands opposite which once it is a cube.

Cutting

A solid with a flat plane through it, and the outline the plane leaves on the newly exposed face.

How the set is put together

Items
24 — eight turning, eight folding, eight cutting, shuffled together
Clock
8 minutes over the whole set, not per item
Choices
four per item, answered with 1 to 4 or by tapping
Order
never more than two of a kind in a row
Scoring
three separate counts, each with a 95% Wilson interval and an exact binomial test against one in four
Not scored
practice items, items the clock ran out on, and items withdrawn when the tab went away

The practice items are outside the 8 minutes, name the right answer afterwards, and reach none of the three sub-scores. The clock starts on the first scored item and stops whenever this tab is not in front, so time spent in another window is not charged to you.

How to take the spatial reasoning test

One budget, three kinds of item, four options each.

  1. Take the three worked items first

    One turning item, one folding item and one cutting item run before anything counts, each with the reasoning shown afterwards. They exist because the folding item is the one people misread: it asks which face ends up opposite a marked face once the net is a cube, which is a relation the flat drawing does not show anywhere. Two minutes here saves the eight that follow.

  2. Watch the budget rather than the item

    No item has a limit of its own. One countdown runs across the whole set, so you can spend four minutes on a cube net if you decide it is worth four minutes — and the seven items after it will have to share what is left. Skipping is not offered, because an item left blank and an item answered by eliminating two options are different events and the run needs to be able to tell them apart.

  3. Read the three columns, not the total

    The result gives turning, folding and cutting as separate counts out of eight with a 95 percent interval on each, and then says in one line whether your best and worst intervals actually overlap. If they do, the gap between them is noise and the honest reading is that all three came out much the same. There is no combined score anywhere on the page because adding the three would undo the only useful thing the run does.

Technical specifications

Items24 scored — eight turning, eight folding, eight cutting — shuffled so the three kinds never arrive in a block, plus three worked practice items that reach no figure
BudgetEight minutes across the whole set, about 20 seconds an item if spent evenly. The countdown stops when this tab goes behind something, so the budget measures time you were actually looking at it
Turning itemsChiral pentominoes — F, L, N, P, Y and Z, the six five-square shapes that differ from their own mirror image. Each option holds five squares, so counting squares eliminates nothing and the shape has to be turned
Folding itemsA net of six squares folded into a cube by rolling a cube across it in software, one square at a time. Only the arrangements where all six faces land distinctly are used, which is what makes a net a net
Cutting itemsFourteen solid-and-plane cases across six solids — cone, cylinder, square pyramid, cube, sphere and triangular prism — each cut at an angle where the outline is not the one the solid suggests
ResponseKeys 1 to 4 on the number row or the keypad, and four on-screen buttons that do the same thing. Four options put chance at 25 percent, so two of eight right is what guessing produces
What voids an itemLosing the foreground voids the question on screen, and so does an answer arriving between two questions. Voided items are listed by name in the result instead of being counted as errors
What leaves the pageNothing. The items are built in the tab from a seed drawn when you press start, and the seed is printed with the result so you can say which run a score came from

Frequently asked questions

Why three sub-scores instead of one spatial reasoning score?

Because the three operations separate in people, not just in theory. Decades of factor-analytic work on spatial ability keep returning a small number of distinguishable factors rather than one, and the everyday version of that finding is familiar: the person who can pack a car boot perfectly is not reliably the person who can tell you what a cut through a cone looks like. A single total made from eight items of each would be an average over three things that disagree, and averaging over disagreement is how a number stops meaning anything.

What is a good result on 24 items?

There is no threshold on this page and there is a reason it is missing. Eight items is a small sample, and a count out of eight carries an interval roughly 30 percentage points wide even when everything goes well — which is why the interval is printed beside each sub-score rather than a label. What the numbers can support is a comparison inside your own run: if your best and worst sub-scores have intervals that do not overlap, that difference is real, and if they overlap, it is not.

Can I stop the clock or take longer than eight minutes?

No, but the clock is fairer than a fixed per-item limit. Eight minutes is the whole budget and you decide how to spend it, so a hard cube net can be given as long as you think it is worth. The countdown pauses only when this tab loses the foreground, and that pause is not a break you can use — the item on screen when the tab went away is voided rather than resumed, because you would come back to a question you had already been looking at.

Why are the shapes drawn by the page rather than photographed?

Because a picture cannot be checked and a construction can. Every figure here is the output of the same arithmetic that produced the answer: the cube is folded by rolling a cube across the net and recording where each face lands, and the cross-section outline comes from the plane and the solid rather than from a draughtsman's idea of it. That makes a wrong item impossible in a way that a library of images never is, and it means a second run genuinely contains different items instead of the same twenty-four in a new order.

Is this the spatial section of a job aptitude test?

It is the same family of item, put to a different use. Employment tests in this area are supervised, timed to the item, and scored against a norm group the publisher collected — none of which applies here. What does carry across is the practice: the item types are the ones those batteries use, so working through these and reading the explanations tells you which of the three you should spend your preparation on. The score itself does not transfer and this page never suggests it does.

Which of the spatial tests here should I take after this one?

Whichever sub-score came out lowest, because the specialist pages measure the same operation with far more items and a much sharper report. Eight items can tell you where to look and cannot tell you much more than that. The composite exists to point; the pages it points at are the ones that measure.

Does this get easier with practice, and does that spoil the measurement?

It gets easier, and it does not spoil anything as long as you know which run is which. Spatial tasks improve with practice more readily than most cognitive tasks, and the improvement is real rather than an artifact — people genuinely get better at folding nets after folding a few dozen. That means a second run is not a re-measurement of the same thing, and a run taken cold is the only one comparable to somebody else's cold run.

What splits when spatial ability splits

The three item types here are not three difficulty levels of one skill. Turning asks you to keep a rigid shape rigid while its orientation changes, which is why the wrong options are mirror images rather than variations — a mirror image cannot be reached by any rotation, so the item has an exact answer and no near-misses. Folding asks something structurally different: the net has to stop being flat, and the relation the question wants, which face ends up opposite which, is not drawn on the net at any point. Cutting asks for neither. It asks what a plane leaves behind, and its characteristic error is answering with the shape of the solid instead of the shape of the section — the cone that produces an ellipse, the cube whose diagonal cut produces a hexagon.

The best-quantified of the three is turning, and the figure that quantifies it is a rate rather than a score: ≈ 17 ms per degree, from Shepard & Metzler (1971), Mental rotation of three-dimensional objects, Science. That number is not printed anywhere in this run, deliberately, because reading a rate off eight untimed items would be arithmetic pretending to be measurement. It is measured properly, with the response time plotted against the angle, on the mental rotation test. The folding items here are a two-dimensional cousin of the punched-hole task on the paper folding test, which runs sixteen items and opens the sheet back out stage by stage after each practice item so a wrong reconstruction is visible rather than merely marked. If the sub-score that came out low was neither of those, the construction version of the same ability is on the block design test, where the pattern has to be built rather than recognized.

Two abilities that people expect to find here are somewhere else on purpose. Knowing where things are relative to you while you move among them — is perspective taking, and it is measured in degrees of pointing error on the spatial orientation test rather than in items right; the applied form of the same thing, with bearings and grid references, is on the map reading test. And spatial reasoning is one input to a general-ability score rather than the score itself, which is worth remembering before reading a low sub-score as a verdict on anything broader — the pages that attempt breadth, such as the Mensa-style IQ test or the speed-of-processing sweep on the brain age test, sample several abilities precisely because no one of them stands in for the rest.

A reaction time here is the interval between the frame that painted the stimulus and the timestamp the browser attached to your key, both read from the same monotonic clock. What neither can see is the display pipeline behind it, so on a 60 Hz screen roughly 16 ms of every figure below is the machine rather than you. That is the timing floor: two numbers closer together than that are the same number, and this page reports no precision it cannot support.

That floor matters less here than on the reaction-time pages, because the figures this run leads with are counts out of eight rather than durations. Where it does apply is the median time per sub-scale, which is printed to help you see whether a low count came from rushing or from running out of budget, and which should not be read to a precision finer than a frame.

This is a measurement exercise, not a clinical assessment. It reports what you did on this page against a stated reference and nothing more — it cannot establish a visuospatial impairment or a learning difficulty. Only a qualified professional, working with more than a browser, can make that judgment.

A run of eight items per sub-scale cannot separate a genuine weakness from an unfamiliar interface, an interrupted afternoon or a screen too small to see the net on. If spatial difficulty is affecting driving, navigation or work, the assessment for that is conducted by a person, not by a page.

Where the 24 items come from and where the answers go

Every number on this page is worked out by JavaScript running in the tab you are reading it in. Your answers, your reaction times and your score are never uploaded, logged or kept — which is also why the test carries on working after you disconnect from the network, and why nothing here can be held back behind an email address.

The seed is drawn in your browser when you press start, and the pentominoes, nets and cross-sections are constructed from it on the spot. No item text, no answer and no timing leaves the tab at any point in a run, and closing the tab is what deletes the result — there is nothing else holding a copy.