Skip to content
AbilityBench

Insight

Problem solving test that watches how the answer arrived

Eight problems, four you can make visible progress on and four you cannot, with no label saying which is which — free, no account, up to 90 seconds each. Every 15 seconds the page asks how close you feel on a scale of 1 to 7, and at the moment you answer it asks how sure you are. The number it reports is the step between those two, because on a problem you were genuinely working through it is small and on a problem that arrived whole it is large.

  • 100% free
  • No signup
  • 8 problems
  • Warmth rating every 15 s
  • Answers shown, not sold

Eight problems, four of each kind, up to 90 seconds each. Every 15 seconds the page asks how close you feel to the answer on a scale of 1 to 7, and when you finally answer it asks how sure you are. The point of the run is the difference between those two numbers, which behaves completely differently on the two kinds of problem.

What the two kinds of problem are

Analytic
A problem you can be visibly halfway through. The equation is set up, the unknown is nearly isolated, and if somebody stopped you at 40 seconds you could say how far you had got.
Insight
A problem where halfway is not a place. You are working on the wrong picture of it until the picture changes, at which point the answer is already there. Stopped at 40 seconds you would have nothing to show, and 5 seconds later you would have all of it.

You are not told which kind any problem is. Being told would wreck the measurement, because expecting a flash changes how long you are willing to sit with something.

The warm-up problem has no clock on it at all and does not reach any figure — it is there so the rating rhythm is familiar before anything is being measured. Answers are numbers; units and currency symbols are ignored.

How to take the problem solving test

Two kinds of problem, one rating scale, and a step at the end of each.

  1. Work on the problem and rate how close you feel

    A problem appears with a 90-second clock. Every 15 seconds a rating falls due: 1 means no idea where this is going, 7 means one line away from the answer. The rating is not a formality — it is the measurement, which is why the answer box locks until you have given the one that is waiting. Answers are numbers, and units, currency symbols and commas are all ignored.

  2. Say how sure you are, then submit

    Before the answer counts you pick a confidence on the same 1-to-7 scale. Do it honestly rather than modestly: the page subtracts your last warmth rating from this number, and a habit of always saying 4 flattens the one quantity the run is about. If the 90 seconds run out first, the problem is recorded as unsolved rather than wrong, and the run moves on.

  3. Read the two shapes, then tick anything you had already met

    The result puts the analytic and insight problems side by side: accuracy with intervals, warmth gained per 15 seconds, the step at the answer, and whether your confidence separated your right answers from your wrong ones. Then it lists the problems with their solutions and a checkbox for each — tick the ones you had seen before and every figure recomputes without them, because a remembered answer has a warmth curve belonging to memory rather than to problem solving.

Technical specifications

Items8 scored problems drawn from a bank of 10 — 4 analytic and 4 insight — plus one optional untimed warm-up that is never one of the eight
Time per problem90 seconds, after which the problem closes itself and is recorded as unsolved. The warm-up has no limit at all
Rating scheduleOne warmth rating every 15 seconds, so a problem carries up to 5 of them before the window closes. The interval is the one Metcalfe and Wiebe used, kept unchanged so the shapes are comparable to theirs
ScalesWarmth and confidence both run 1 to 7. They are the same scale on purpose: the headline figure is a subtraction between them, and two differently sized scales would make that subtraction meaningless
Answer checkingEvery answer is a single number, compared after commas, currency symbols, units and surrounding words are stripped. There are no multiple-choice options anywhere, because four options let an insight problem be solved by testing them
What the run reportsAccuracy per kind with 95% Wilson intervals, mean warmth gained per 15-second interval, the step from last warmth to final confidence, mean time per kind, and the confidence gap between right and wrong answers
Familiar itemsRecorded after the run by your own tick, and excluded from every figure when ticked. None of the three widely circulated cognitive-reflection items is in the bank at all
Reference figuresTwo shapes rather than a scale: Metcalfe & Wiebe (1987), Intuition in insight and noninsight problem solving, Memory & Cognition. No percentile is produced from 8 problems, and none of the items has a published solution rate this page would stake a number on

Frequently asked questions

What actually makes a problem an insight problem?

Being partway through it is not a state that exists. On an analytic problem the work accumulates — the equation is set up, the unknown is nearly isolated, and if somebody stopped you at 40 seconds you could show them where you had got to. On an insight problem you are working with the wrong picture of the situation until the picture changes, and the moment it changes the answer is already there. That is why the two kinds produce different rating curves rather than merely different difficulties, and why a problem can be easy and still be an insight problem.

Why does the page keep interrupting me for a rating?

Because the ratings are the data and the answer is almost incidental. Two people can both solve six of eight and have solved them in completely different ways, and the only visible trace of that difference is what they were saying about their own progress while it happened. Metcalfe and Wiebe's 1987 experiment is built entirely on those interruptions, at exactly this 15-second spacing, and a version that collected the answers without the ratings would be a quiz. It is also why the submit button locks while a rating is outstanding rather than letting you skip it.

I had seen one of these before. Does that ruin it?

It ruins that problem, which is why the page asks and then takes it out of the figures. A remembered answer is retrieved rather than found: the warmth curve belongs to memory search, the confidence at the answer is high from the first rating onwards, and the step this page measures collapses to nothing. The three items from the classic cognitive-reflection test are excluded from the bank entirely for the same reason — the bat-and-ball question is among the most circulated puzzles on the internet, and a page that used it would be measuring internet exposure.

Why numbers instead of multiple choice?

Because four options turn an insight problem into an analytic one. Given a list, you can test each candidate against the wording and arrive at the answer by elimination, which is a procedure — and procedures are precisely what the insight items are constructed to deny you. A typed number also has a side benefit for the ratings: with nothing to check against, a warmth rating of 6 has to be about your own sense of the problem rather than about how many options you have ruled out.

My confidence tracked both kinds equally. What does that mean?

Most often it means the ratings were given quickly rather than considered, and sometimes it means the insight problems were solved analytically. Both are worth knowing. The published pattern is that a mid-problem rating predicts success on analytic problems and predicts very little on insight ones, so a run where both predict equally well suggests you were treating all eight the same way — which is a real strategy and produces a real, flat result. Four problems per kind also leaves plenty of room for the pattern to disappear into noise, and the intervals printed beside every proportion are how wide that room is.

Is being bad at the insight items a problem?

No, and the reason is that insight items are famously unrelated to the things people usually mean by problem-solving ability. Performance on them correlates weakly with performance on analytic problems, they are strongly affected by the exact wording, and the same person can crack one and be defeated by another of identical structure. Nothing on this page is a measure of general ability, and eight items could not be one even if the items were chosen for it. What the run can show you is how one kind of item behaved against the other in your own hands, which is a much smaller and much more defensible claim.

What happens to the answers I typed?

They stay in this tab and are compared with a number, once, and then they exist only in the results panel until you close it. Nothing is submitted anywhere — there is no request to any server during a run, which is also why the test keeps working with the network off. The copy button writes a plain-text summary of the figures to your own clipboard and includes no problem text and no answers.

Two kinds of problem, and the experiment that separated them

In 1987 Metcalfe and Wiebe gave people two sorts of problem and interrupted them every fifteen seconds to ask how warm they were getting. On algebra problems the ratings climbed steadily, and a rating taken halfway through predicted whether the problem would be solved. On insight problems the ratings stayed flat and then jumped — and a rating taken halfway through predicted almost nothing, including on the problems that were about to be solved seconds later. The finding is not that insight problems are harder. It is that people have accurate access to their progress on one kind and almost none on the other, which means “feeling stuck” is informative in one case and worthless in the other. Metcalfe & Wiebe (1987), Intuition in insight and noninsight problem solving, Memory & Cognition.

That asymmetry is the practical content of this page, and it has a use outside a psychology experiment. If you are working on something where progress accumulates, a persistent sense of being nowhere is evidence and you should probably change approach. If you are working on something whose difficulty is that you have the wrong picture of it, the same feeling is not evidence of anything at all, and the correct response is to keep going or to leave it and come back — the two standard pieces of advice about being stuck are each right about one kind of problem and wrong about the other. What no amount of introspection tells you is which kind you are holding, which is why this run does not tell you either: labeling the problems would remove the only interesting thing about the ratings.

Where this version departs from the original is worth being explicit about. The problems are written for this page rather than taken from any published set, so the difficulty levels are not calibrated against anybody; the insight items use classic structures with new wordings, which is why the page asks afterwards which ones you had already met. Eight problems is also a small number: four per kind puts a 95% interval on an accuracy roughly forty points wide, and the page prints those intervals rather than a percentage on its own for that reason. The one figure that survives a short run is the within-you comparison — your own step on one kind against your own step on the other. If you want the neighboring measurements instead, the exercise where the sudden part is a plan rather than an answer is the Tower of London test, the puzzle where the structure is fully known and only your route through it varies is Tower of Hanoi, and the short battery that samples inhibition, switching and planning in one run is the executive function assessment. For what survives when a competing stream has to be ignored, the selective attention test and the concentration test are the two to run.

Why the exclusion checkbox is not optional politeness

Cognitive Reflection Test mean score: 1.24 of 3 on the original three items, among students at us universities.

Frederick (2005), Cognitive Reflection and Decision Making, Journal of Economic Perspectives

The sample was university students, so this is not a population average and a visitor scoring above it has not beaten the general public. The larger problem is exposure: these three items are among the most circulated questions on the internet, and someone who already knows the ball costs five cents is not demonstrating reflection, which is why a score is worth reading only alongside how many of the items you had met before.

On precision: the clock on each problem starts at the frame that painted it and every rating is stamped against that same instant, both read from one monotonic clock. Nothing here is reported finer than a second, because a fifteen-second rating interval and a forty-second solve have no use for a millisecond — and a page that printed one would be claiming an accuracy the measurement does not contain.

This is a measurement exercise, not a clinical assessment. It reports what you did on this page against a stated reference and nothing more — it cannot establish a learning difficulty, ADHD or any thinking disorder. Only a qualified professional, working with more than a browser, can make that judgment.

Where your answers and ratings live

Every number on this page is worked out by JavaScript running in the tab you are reading it in. Your answers, your reaction times and your score are never uploaded, logged or kept — which is also why the test carries on working after you disconnect from the network, and why nothing here can be held back behind an email address.

The problems, the answers you type, the ratings you give and every figure computed from them stay in this tab. Nothing is written to storage, so reloading loses a finished run, and the copy button carries the summary figures only — no problem text, no answers and no rating-by-rating list.