Rule diagnosis
Matrix reasoning test that tells you which rule you missed
Fifteen 3x3 grids with one cell missing, five built on a single rule, five on two and five on three, shuffled together so difficulty is not confounded with fatigue — free, no account, no time limit. Every wrong tile is the right cell with exactly one property altered, so the report names the property you got wrong and lets you walk back through each missed item with your tile beside the correct one.
- 100% free
- No signup
- 15 items, 5 per rule count
- Every miss reviewable
- Error split by property
15 matrices, five carrying one rule, five carrying two and five carrying three, shuffled together rather than ordered easy to hard. The shuffle is the point: an ascending set makes the hard items also the late items, and then a drop in accuracy could be difficulty or could be that you had been at it for fifteen minutes.
One worked out, before you start
This grid is a real generated item with its answer shown. The rules that produced it are listed underneath. Nothing here is scored.
Missing cell
Rules in it
- Figure addition or subtraction on which lattice positions are occupied
- Constant in a row on which shape
Eight tiles per item, no clock, no going back. Every wrong tile is the correct cell with exactly one thing altered, which is what makes the report at the end possible: the tile you pick says which rule you were working from, not merely that you were wrong.
How to read a per-rule breakdown
Three outputs, in the order they are worth looking at.
Start with the slope, not the total
The first panel is accuracy at one, two and three rules per item, with five items behind each figure. A flat profile and a falling one mean different things: falling means the rules were available to you individually and the load of holding several at once was what cost you, which is the pattern the format was designed to produce. Each figure carries a 95% Wilson interval, and on five items those intervals are wide enough to overlap, so read the direction rather than the gap.
Find the rule that keeps appearing in the miss column
The second table counts every item that contained each of the five rules and how many of those you solved. Items carrying three rules count under all three, so the rows deliberately sum past fifteen. A rule you lost most of the time is a rule to go and read the statement of; distribution of two values is the usual candidate, because its third value is an absence and an absence is not a thing the eye reports.
Walk through the misses one at a time
The review pane shows each item you got wrong with the grid, the tile you chose and the tile that was correct, and states what your tile altered — a property that carried a rule, a property that was the same in all nine cells, or a straight copy of a cell already on screen. That last category is the informative one: choosing a cell you have already seen means the eye found a match before the rules were checked.
Technical specifications
| Items and their mix | 15 scored items: exactly 5 carrying one rule, 5 carrying two and 5 carrying three. Equal cells are the reason the three accuracy figures can be compared at all |
|---|---|
| Presentation order | Interleaved, with no more than two items of the same rule count in a row. The sibling page presents its items easiest-first because the format is progressive by design; this one refuses to, because an ascending order makes the hard items also the tired items |
| Rule coverage | The set is regenerated until each of the five rules appears in at least 2 of the 15 items, so no row of the rule table stands on a single item |
| Distractor construction | 7 per item. Each is the correct cell with exactly one property changed — shape, shading, count, size, orientation, the extra mark or the occupied positions — or a copy of the cell to the left, above or at the center. The change is recorded, which is what makes the error report possible |
| What the review shows | Every missed item, with the full grid, your tile, the correct tile and the list of rules that generated it. Correct items are not replayed: the answer to a solved item is the tile you already chose |
| Timing | Untimed per item; each duration is measured from the frame that painted the grid to the timestamp on your click and reported as a median per rule count. Rising medians across the three levels are the effort of holding more rules at once |
| Reference scale named, not applied | Wechsler subtest scaled score: mean 10 points, SD 3 — Wechsler subtest scaled scores are constructed with a mean of 10 and a standard deviation of 3 — The composite indexes built from them use a mean of 100 and a standard deviation of 15, which is why the two numbers in a report look so different.. Printed to explain what this page cannot return; nothing here is converted onto it |
| What leaves the page | Nothing. The review, the tables and the copied summary are all built from an array of records held in this tab, and the summary carries counts rather than your per-item answers |
Frequently asked questions
Is matrix reasoning the same thing as a whole matrix test?
No — the phrase usually names a subtest inside a larger battery rather than a test on its own. In a clinical or hiring battery, matrix items are one of several tasks whose scores combine into an index, and the subtest is short precisely because it is not carrying the whole estimate on its own. That is worth knowing if you have a report in front of you: the matrix line in it was never meant to be read alone, and neither is a fifteen-item run here.
Why does this page shuffle the difficulty instead of building up to it?
Because the output is a comparison between difficulty levels, and an ascending order would put every hard item at the end of the session. Whatever the last five items measure, part of it is that they came last: attention has drifted, the novelty has gone, and the visitor has decided how much effort this is worth. Interleaving costs the run the teaching effect a progressive order gives, which is a real loss and is why the sibling page keeps it.
I keep choosing tiles that were already somewhere in the grid. Why?
That is the repetition error and it is generated deliberately, because it is the commonest wrong answer in this format. It happens when one property has been solved and the rest of the tile is matched by recognition rather than by rule: a cell you have looked at four times already feels correct, and feeling correct is exactly what a distractor is built to exploit. The fix is mechanical — before choosing, check the candidate against every rule you found, not the one you found first.
Fifteen items is not very many. What can it actually support?
It supports the error structure and not much about your level. Fifteen items give a 95% interval on overall accuracy roughly 45 percentage points wide, and five items per level give intervals wide enough that the three usually overlap — so the direction of the slope is readable and its size is not. What does not depend on sample size is the composition of your mistakes: if four of five misses altered the shading, that is a fact about the four items regardless of how many there were.
My time per item doubled on the three-rule items. Is that bad?
It is the expected pattern and it is the more interesting half of the result. Rules have to be found one at a time and then held while the candidates are checked, so an item with three of them costs roughly three searches plus the bookkeeping of keeping the first two available while the third is worked on. A profile where time rises steeply and accuracy holds is a different thing from one where time stays flat and accuracy falls: the first bought its accuracy with effort, and the second ran out of places to put the rules.
Can I get a scaled score or an index score out of this?
No. A scaled score is a rank against an age-banded standardization sample, collected under supervision and published only to licensed purchasers of the test kit, so no browser page has the sample the conversion needs. What this page returns instead is a count out of fifteen, three level accuracies with intervals, a rule table and an error breakdown — all of which are statements about the fifteen items you were shown.
Why does the review only show the items I got wrong?
Because a correct item has nothing to explain: the tile you chose is the answer, and replaying it would pad the review to fifteen screens with fourteen of them saying you were right. The rules that generated every item are in the copied summary if you want the full inventory. If you want the items themselves back, the seed in that summary regenerates the exact fifteen.
Where the phrase comes from, and what an error report can honestly claim
“Matrix reasoning” entered common use as the name of a subtest, not a test. In a full cognitive battery it is one of several tasks feeding a perceptual or fluid reasoning index, it runs to a couple of dozen items, and its own score is reported on a scale built with a mean of 10 and a standard deviation of 3 — a metric that means nothing without the age-banded sample behind it. Those samples are the reason this page returns none: The norm tables are the product. They are copyrighted, sold with the test kit under a qualification requirement, and reproducing them would be both an infringement and an invitation to treat a browser imitation as a Wechsler score.
What a generated set can do honestly is describe the shape of your errors, and that needs one design decision that most online matrix tests skip. A distractor drawn at random is nearly always eliminable at a glance and teaches nothing when it is chosen. Every distractor here is instead the correct cell with exactly one property altered, and the alteration is recorded — so a wrong tile is a small piece of evidence about which rule you were running. Choosing a tile that breaks a property carrying a rule means the rule was not found. Choosing one that breaks a property identical in all nine cells means the tile was picked before it was checked. Choosing a copy of a cell already on screen is the repetition error, and it is the one that separates a visitor working by rule from a visitor working by resemblance. Those three categories are the report, and none of them needs a norm to be true.
The second output is accuracy against the number of rules per item, which is where this page and the instrument version of the same generator genuinely part company: that page orders items easiest-first because the published format teaches through its ordering, and this one shuffles them because a slope measured in presentation order is partly a fatigue curve. If you want the same items under an employer’s clock, that is the abstract reasoning test; if you want the rule-finding cost isolated rather than the rule-holding cost, the inductive reasoning test times how long the first solution in a family takes. A battery that leads with matrix items usually also contains a critical-reading section, which is the shape the Watson-Glaser practice test takes, and an orthography section, which is the spelling test.
The scale this page names but does not use
Wechsler subtest scaled score: mean 10 points, SD 3.
Wechsler subtest scaled scores are constructed with a mean of 10 and a standard deviation of 3 — The composite indexes built from them use a mean of 100 and a standard deviation of 15, which is why the two numbers in a report look so different.
This is the shape of the scale, not a figure anybody was measured at. A scaled score exists only relative to an age-banded standardization sample administered under supervision, so nothing on this page can produce one, and a raw count out of fifteen cannot be converted into one afterwards.
A reaction time here is the interval between the frame that painted the stimulus and the timestamp the browser attached to your key, both read from the same monotonic clock. What neither can see is the display pipeline behind it, so on a 60 Hz screen roughly 16 ms of every figure below is the machine rather than you. That is the timing floor: two numbers closer together than that are the same number, and this page reports no precision it cannot support.
The medians this page prints are in seconds and the floor under them is one frame, so it changes nothing at the precision shown. It is stated because the same clock and the same rule produce the millisecond figures elsewhere on this site, and a page that mentioned its error bars only when they were small would be advertising rather than reporting.
This is a measurement exercise, not a clinical assessment. It reports what you did on this page against a stated reference and nothing more — it cannot establish dyslexia, an attention disorder, or any developmental or intellectual condition. Only a qualified professional, working with more than a browser, can make that judgment.
Where the fifteen items and your fifteen answers stay
Every number on this page is worked out by JavaScript running in the tab you are reading it in. Your answers, your reaction times and your score are never uploaded, logged or kept — which is also why the test carries on working after you disconnect from the network, and why nothing here can be held back behind an email address.
The review pane is the reason this page holds more in memory than most: every item, its options and the tile you chose stay in a JavaScript array so they can be shown back to you. That array is never serialized, never stored and never transmitted, which is also why reloading loses the review and why the copy button exists.