Rule discovery and rule change
Card sorting by a rule that changes without telling you
Put each card on one of four piles; the only thing you are told is right or wrong, and once ten sorts in a row come back right the rule silently becomes a different one. Free, 64 cards, four to seven minutes, and the counts appear the moment the deck runs out with nothing held back. What the run measures is not how quickly you find a rule but how expensively you let go of one, which is the count printed largest.
- 100% free
- No signup
- 64-card deck
- Rule changes silently
- Raw counts, no percentile
A card appears; you put it on one of four piles. Nobody tells you what makes a pile the right one — the only information you get is the word right or wrong after each card, and once you have found the rule it changes, without any announcement, and keeps changing. Up to 64 cards, usually four to seven minutes.
The procedure, in full
- Piles
- one red triangle, two green stars, three yellow crosses, four blue circles
- Deck
- all 64 combinations of the four colors, four shapes and counts one to four, shuffled fresh
- Feedback
- the word right or wrong for 480 ms, then a gap of 140 to 260 ms drawn at random so the next card cannot be timed
- Change
- 10 correct sorts in a row complete a category and the rule silently becomes the next one
- Order
- drawn at random for this run, not the published sequence
- Spill
- a click inside 100 ms of a card appearing belongs to the previous card; that card leaves the deck uncounted
- Ends
- at the 64th card or the 6th category, whichever arrives first
Which of the three came up first, and in what order the rest followed, is held back until the deck is finished.
The paradigm: Grant & Berg (1948), A behavioral analysis of degree of reinforcement and ease of shifting to new responses in a Weigl-type card-sorting problem, Journal of Experimental Psychology
The practice cards are announced as a color sort and are drawn separately, so they cost you none of the 64 and tell you nothing about which rule the deck opens on.
How to run the card sort
Guess, read the feedback, and keep the feedback in mind for longer than one card.
Sort the first card by guessing
Four piles sit above the card you have been dealt: one red triangle, two green stars, three yellow crosses, four blue circles. Nothing about the first card tells you which pile is correct, because the rule has not been revealed and never will be. Click a pile, or press its number key, and read what comes back.
Turn right and wrong into a rule
There are only three things a pile can have in common with your card: its color, its shape, or how many symbols are on it. Two or three consistent answers are usually enough to tell which of the three is being rewarded. There is no clock on a card and no penalty for thinking, so the useful move after a wrong answer is to work out what your choice ruled out.
Notice the moment the rule stops working
After ten correct sorts in a row the rule changes to one of the other two, with no announcement, no pause and no change in how the cards look. The first card of the new rule will come back wrong however well you had been doing. What the run counts is how many cards you spend defending the old rule before you try a different one, and it separates that from errors that are simply guesses.
Technical specifications
| Deck | All 64 combinations of four colors, four shapes and one to four symbols, shuffled fresh each run. Because the deck is the whole space rather than a sample of it, every card has exactly one correct pile under each of the three possible rules |
|---|---|
| Piles | Four, fixed and never reordered: one red triangle, two green stars, three yellow crosses, four blue circles. Number keys 1 to 4 and the numeric keypad answer as well as a tap |
| Rule change | Ten correct sorts in a row complete a category and move the rule on. Six categories is the ceiling, which needs 60 of the 64 cards sorted correctly with no run ever broken |
| Rule order | Color, shape and count are put in a random order for each run and then cycled, so consecutive categories always differ. The published order is fixed and widely reproduced, which would hand ten free sorts to anyone who had read it five minutes earlier |
| Feedback | The word right or wrong for 480 ms, then a gap drawn at random between 140 and 260 ms. No running score, no streak counter and no signal that a category has just ended, and a gap you cannot time means the next card cannot be beaten to |
| Pace | No time limit on a card and no clock on screen. A click inside 100 ms of a card appearing is treated as spill-over from the card before and that card leaves the deck uncounted. The median seconds per card is reported afterwards and feeds none of the counts, because a deliberate sorter and a fast one reach the same categories |
| Reported | Perseverative errors, other errors, categories completed, cards to the first category, runs of five broken before ten, and the correct sorts that also fitted the rule you had just left |
| What leaves the page | Nothing unless you press copy, which puts the counts, the rule order this run used and the run seed on your own clipboard as plain text |
Frequently asked questions
Why is a wrong sort just after a rule change counted differently from any other?
Because it is the one error that carries information about what you did rather than about what you had not yet learned. Every wrong sort before you have found a rule is a guess, and guesses are cheap and uninformative. A wrong sort straight after a change is a card placed by the rule that the feedback has already stopped rewarding — you had the right idea and the evidence for it was ten cards deep, so abandoning it costs something. Counting the two together produces one number in which a person who was guessing and a person who could not let go look identical, and separating them is the entire reason this paradigm exists.
How many cards should it take to find the first rule?
There is no defensible number to give you, and this page will not print one. The first card of a run is a pure guess with a one-in-four chance, the second and third narrow it, and from there it depends on which cards the shuffle happens to deal — a run of cards whose color and shape point at the same pile tells you much less than one where they disagree. What the run does report is how many cards your first category actually took, which is comparable against your own second run and against nothing else.
The rule changed and I got four wrong in a row. Is four bad?
No, and the shape of the sequence matters more than the count. One or two wrong cards after a switch is close to unavoidable: you cannot know a rule has changed until a card that should have been right comes back wrong, and the second card is where the search begins. What the perseverative count is looking for is the difference between two cards of resistance and eight, and between an error that repeats the old rule and one that tries something new and misses. Both appear separately in the table under the result.
Does being colorblind make this test unfair?
It would, so there is a switch that removes the problem without changing the measurement. Turning on the attribute labels prints each card's color, shape and count underneath it in words, on the four piles as well as on the card you are sorting. That gives nothing away, because the difficulty here has never been in perceiving the attributes — it is in working out which of them is being rewarded and in noticing when that stops being true. Screen readers get the same description from every card whether the labels are showing or not.
Why does the page not tell me when I have completed a category?
Because being told a rule has changed removes the thing being measured. The whole cost this task quantifies is the interval between a rule ceasing to work and you accepting that it has, and an announcement collapses that interval to zero. It is also why there is no streak counter on screen: watching a run climb toward ten tells you exactly when the change is about to arrive, and you would then be waiting for it rather than being caught by it. The category structure is revealed in full once the deck is finished.
Is this the published Wisconsin card sorting test?
No, and the differences are deliberate rather than incidental. This is the generic sorting paradigm the published instrument was built on, run in a browser: the deck order is generated, the rule order is randomized rather than fixed, and perseveration is counted by a rule stated in full on the results panel instead of by the scoring manual's more elaborate one. That manual and its normative tables are a commercial clinical product and are not reproduced anywhere here, so no count on this page can be looked up in them and none is presented as a score on that test.
Why 64 cards rather than the 128 the full version uses?
Because the second 64 are the same deck again and mostly measure stamina. The structure this task is built around — find a rule, hold it, lose it, find the next one — is fully present in one pass, and a run of 64 reaches four categories for most people who reach any. A 128-card version in a browser, with no examiner in the room and nothing at stake, is where people start clicking to get to the end, and a run somebody stopped attending to halfway through produces perseverative counts that mean nothing.
What a card sort measures, and the claim this page will not make
The paradigm is a small piece of experimental design that has outlived almost everything around it. Grant & Berg (1948), A behavioral analysis of degree of reinforcement and ease of shifting to new responses in a Weigl-type card-sorting problem, Journal of Experimental Psychology set out the arrangement that is still in use: a stimulus card, several piles it could belong to on more than one attribute, and a rule of reinforcement that the participant has to infer and that the experimenter then changes. The asymmetry it exposes is the interesting part. Learning the rule is easy — almost everybody finds it inside a handful of cards — while giving it up is not, because by the time the rule changes it has been confirmed ten times and the card contradicting it has been seen once. The count that describes that resistance is the perseverative error, and it is the reason this task survived when most of its contemporaries did not.
The claim this page refuses is the familiar one: that a card sort is a frontal-lobe examination. It became attached to that idea because early lesion work used it, and the attachment stuck long after the specificity had been argued away — high error counts turn up with poor sleep, with anxiety, with an unfamiliar interface, and with simply not caring very much about a website. A browser version has none of the controls that make a clinical administration interpretable: no examiner watching whether you understood the instructions, no fixed physical deck, no record of what you were doing ten minutes earlier. So the output here is a description of one run and stops there. If you want the same executive machinery approached from a direction where the rule is never hidden and the cost is measured in seconds, that is the trail making test, the other half of this pair; if you want it approached through planning a sequence of moves before making any of them, that is the tower of london test.
Three things go wrong in most browser versions of this task, and all three are fixable. The first is keeping the published rule order, which is printed on every encyclopedia page about the test — a visitor who read one gets the first category free. The second is scoring every error after a switch as perseverative, which quietly merges the people who cannot let go with the people who are guessing and produces a number that only tracks how many mistakes were made. The third is the streak counter: showing a run of correct sorts climbing toward the criterion tells the visitor exactly when to expect the change, and a switch you were waiting for costs nothing to make. Conflict measured at the level of a single response rather than a rule belongs to the simon task, and cancelling an action that has already started belongs to the stop signal task. The nearest relative of all is the balloon analogue risk task, where the thing to be inferred from feedback is not which attribute counts but how far you can push before the feedback arrives at all.
A reaction time here is the interval between the frame that painted the stimulus and the timestamp the browser attached to your key, both read from the same monotonic clock. What neither can see is the display pipeline behind it, so on a 60 Hz screen roughly 16 ms of every figure below is the machine rather than you. That is the timing floor: two numbers closer together than that are the same number, and this page reports no precision it cannot support.
That floor is quoted here because the site quotes it wherever milliseconds appear, and on this page it is the one figure it cannot reach: a card takes seconds to think about and a frame is sixteen thousandths of one. The seconds per card are reported to a tenth for that reason, and every count above — categories, errors, perseverations — is a count of cards and has no clock in it at all.
This is a measurement exercise, not a clinical assessment. It reports what you did on this page against a stated reference and nothing more — it cannot establish a frontal-lobe impairment, ADHD, schizophrenia or any other condition. Only a qualified professional, working with more than a browser, can make that judgment.
Where the deck is shuffled
Every number on this page is worked out by JavaScript running in the tab you are reading it in. Your answers, your reaction times and your score are never uploaded, logged or kept — which is also why the test carries on working after you disconnect from the network, and why nothing here can be held back behind an email address.
The shuffle, the rule order and every card you sorted exist only as JavaScript values in this tab, and reloading the page destroys them. The seed that would reproduce the run is printed in the copied summary and is not sent anywhere, so the only copy of a run you want to keep is the one you take yourself.