Arithmetic, taken apart
Mental math test that reports which operation is costing you
Forty-eight sums typed rather than picked, balanced by construction: twelve additions, twelve subtractions, twelve multiplications and twelve divisions, four of each at small, medium and large operands. Free, no signup, no calculator and no scratch pad. Nothing gets harder as you go — a run that ramps cannot tell you whether your division is slow or whether the division items were the big ones — and because half the additions cross a ten and half do not, the result prices carrying in milliseconds with a confidence interval instead of asserting that it is harder.
- 100% free
- No signup
- 48 sums
- 12 per operation
- Carry cost in ms
Forty-eight sums, twelve of each operation, four at each of three operand sizes, typed rather than chosen from a list. The design is flat on purpose: nothing gets harder as you go, because a run that ramps cannot tell you whether your division is slow or whether the division items were simply the big ones.
Half the additions cross a ten and half do not, and the same split holds for the subtractions that borrow. That is what lets the report put a price in milliseconds on carrying, with an interval around it, instead of asserting that carrying is harder.
No item has a deadline and the run has no total clock, so a pause to think is recorded as a long item rather than as a miss. You can stop at any point and score what you have — the cells are cycled, so a short run is still balanced.
How to run the mental arithmetic drill
Type, press Enter, and read the twelve cells afterwards.
Take the four warm-ups if the keyboard is new to you
One sum from each operation, untimed, with the answer shown afterwards. They exist because the first two scored items of any typed task are slower than the rest for reasons that have nothing to do with arithmetic — finding the number row, discovering that Enter submits — and four throwaway items move that cost out of the measurement. None of them reaches any figure in the report.
Answer forty-eight sums with no feedback between them
Type the answer and press Enter, or use the button. Nothing tells you whether you were right, on purpose: knowing you missed the last one changes how carefully you do the next, and this run needs all forty-eight to be comparable. There is no deadline on any item, so a pause to think is recorded as a long item rather than as a miss, and you can stop at any point — the twelve cells are cycled in turn, so a run of twenty-four items is still balanced across all of them.
Read across the rows, not down the columns
The result is a four-by-three table of medians and accuracies. A row whose three cells barely differ is an operation you have automated; a row that doubles from small to large operands is one you are computing rather than recalling, and that difference is the useful output. Two figures sit beside it: a rank correlation between operand size and solving time, and the cost of carrying as a mean difference with a 95% interval, which is the honest way to say whether this run found the effect at all.
Technical specifications
| Items | 48 scored — a four-by-three design of operation against operand size, four items in every cell — plus 4 optional untimed warm-ups that enter no figure |
|---|---|
| Operand ranges | Small is both operands under 10; medium is 11 to 49; large is 51 to 199 for addition and subtraction, two digits by a single digit for multiplication, and an exact division with an answer from 21 to 39 |
| Division | Always exact, never a remainder. A remainder is a second task wearing the same sign, and mixing the two would put the slowest items in one cell and then label that cell division |
| Carry balance | Six of the twelve additions cross a ten and six do not; six of the twelve subtractions borrow and six do not. That split is what makes a carry cost measurable rather than assertable |
| Times | Taken from correct answers only. A wrong answer is a different event, and a fast wrong one would pull a median down while appearing to show the opposite of what happened |
| Carry cost | Reported as a mean difference in milliseconds with a 95% Welch interval, from the crossing items against the clean ones. Where the interval spans zero the page says the run cannot price it |
| Size effect | Reported as a Spearman rank correlation between operand size and solving time, not as a slope in milliseconds per unit: twelve points per operation will not support a slope, and one long item would set it |
| What leaves the page | Nothing. The forty-eight sums come out of one integer that appears beside your score, and the copied breakdown carries medians and intervals with no item-by-item times in it |
Frequently asked questions
Why does this drill refuse to get harder as I go?
Because a ramp destroys every comparison the report is made of. If the multiplication items arrived later than the additions, they were also bigger, and the gap between the two medians is then a difference in difficulty wearing the costume of a difference in operation. The design is therefore flat: twelve cells, cycled in turn, four passes through all of them. The number a ramp produces is genuinely satisfying, and it lives on the sixty-second sprint instead, which is built to ramp and reports none of this.
Why are my times taken only from the answers I got right?
Because a wrong answer is not a slow correct answer, and averaging the two produces a figure that describes neither. Two kinds of error dominate here and they have opposite durations: a fast wrong answer is usually a retrieved fact that was retrieved incorrectly, and a slow wrong answer is usually a multi-step computation that lost a carry halfway. Pooling them with the correct answers moves the median in whichever direction your particular mix of errors happened to fall. The accuracy figure carries the errors instead, cell by cell.
What does it mean if my large-operand column is much slower than my small one?
It means those items are being computed rather than remembered, which is the ordinary finding rather than a deficiency. Small single-digit facts are stored and retrieved; larger ones are worked out by decomposing, rounding and adjusting, and a procedure takes time in proportion to how many steps it has. The direction of that effect has been documented across mental-arithmetic research for decades — Ashcraft (1992), Cognitive arithmetic: A review of data and theory, Cognition is the usual review — and no magnitude from that literature is quoted here, because a figure collected on a laboratory keyboard is not a figure for yours.
Is a carry cost of 300 ms good or bad?
It is neither, and the interval printed beside it is the only part of that figure that can be read on its own. What the number means is that on this run, the additions and subtractions that crossed a ten took that much longer than the ones that did not; whether the effect exists for you at all is answered by whether the interval clears zero. With six items on each side of the split, an interval two hundred milliseconds wide is normal, and comparing your point estimate with somebody else's is comparing two numbers that both have that much slack in them.
Why is there no scratch pad? I can do these on paper.
Because written arithmetic is a different skill with different limits, and mixing the two would make the operand-size column meaningless. On paper, a three-digit addition costs about what a two-digit one costs — the algorithm is the same and the page holds the intermediate digits for you. In your head it does not, because the intermediate result has to be maintained while the next column is processed, and that maintenance is exactly what the large-operand cells are measuring. A drill that allowed paper would be measuring handwriting speed against a clock.
The keyboard makes me slow. Doesn't that spoil the times?
It adds a constant to every cell, which is why the report is built on differences between cells rather than on their absolute values. Typing 84 takes about as long whether the sum was 79 + 5 or 12 × 7, so the comparison across operations survives it and only the raw medians are inflated. Two places where it does not cancel are worth knowing: a three-digit answer takes one more keystroke than a two-digit one, which slightly exaggerates the large-operand column, and answering with the on-screen button adds a journey to it that pressing Enter does not.
Can I compare my medians with anybody else's?
No, and the reason is worth stating rather than hiding behind a missing table: any figure you found published for a browser arithmetic drill would describe the item difficulty that site chose, so two of them are not on one scale and neither is on a scale with you. What is comparable is your own run against your own run — the same form number rebuilds the same forty-eight sums, which is the only strictly fair repeat — and, inside a single run, one of your cells against another.
Three effects that show up in forty-eight sums on one person
Mental arithmetic is not one operation, and the reason a single score for it is useless is that its parts behave differently. Small products and sums are retrieved: the answer to 7 × 8 is looked up rather than worked out, which is why it can be wrong in a way that is fast and confident. Larger ones are computed by decomposing into parts, operating on them and reassembling, which takes time roughly in proportion to the number of steps and fails in a characteristic way — an answer out by exactly ten, which is a dropped carry rather than a misremembered fact. That is the split the operand-size columns are built to expose, and it is why an error out by ten and an error with the right digits in the wrong order are worth separating before deciding to practice anything.
The second effect is the carry itself, and it is the one this page can actually price for you. Crossing a ten forces an intermediate result to be held while the next column is processed, and that holding competes for the same limited capacity the operands are occupying. Because the design puts six crossing and six clean items in the additions and the same in the subtractions, the difference between the two groups is a measurement rather than a claim — and it comes with a Welch interval, because six items a side does not support a confident point estimate and pretending otherwise would be the same sin as printing an invented norm. The third effect is the ordering of the four operations, where the interesting question is not that division is slowest for almost everyone but whether your multiplication-to-division gap is wider than your addition-to-subtraction gap. A wide one usually means the division is being done as an unfamiliar inverse rather than as a recalled fact.
There is no adult population norm for arithmetic done against a clock in a browser, or for spelling a list somebody assembled. Any figure would depend entirely on the item difficulty the page chose, which makes a comparison across sites meaningless. That is the reason nothing on this page is converted into a rank. The comparison that does work is longitudinal and cheap: run the same form number again in a week and the twelve cells are directly comparable. Around this page, the cluster splits by what you want measured. The sixty-second sprint is the same arithmetic with the analysis stripped out and a published difficulty ladder put in, which is the right shape for a number you want to beat and the wrong shape for a comparison between operations. The math aptitude test removes the clock entirely to ask a question this page cannot: whether the difficulty is fluency or reasoning about quantity. The numeracy test puts arithmetic back inside situations with irrelevant figures attached, and the numerical reasoning test hands the calculator back because the employer format it copies always does. If what you are really chasing is how fast you handle simple material rather than how well you calculate, the processing speed test measures that with no arithmetic in it at all, and the go / no-go test measures the opposite ability — stopping a response that is already under way.
A reaction time here is the interval between the frame that painted the stimulus and the timestamp the browser attached to your key, both read from the same monotonic clock. What neither can see is the display pipeline behind it, so on a 60 Hz screen roughly 16 ms of every figure below is the machine rather than you. That is the timing floor: two numbers closer together than that are the same number, and this page reports no precision it cannot support.
On this page that floor matters to exactly one figure. The medians are in seconds and carry it invisibly, but the carry cost is printed in milliseconds, and a point estimate of 40 ms would be inside the same order of magnitude as one frame of your display. That is why it never appears without its interval: an effect smaller than the machinery measuring it is reported as an interval that includes zero, which is the true statement.
This is a measurement exercise, not a clinical assessment. It reports what you did on this page against a stated reference and nothing more — it cannot establish dyscalculia or an arithmetic learning difficulty. Only a qualified professional, working with more than a browser, can make that judgment.
Where forty-eight solving times are worked out
Every number on this page is worked out by JavaScript running in the tab you are reading it in. Your answers, your reaction times and your score are never uploaded, logged or kept — which is also why the test carries on working after you disconnect from the network, and why nothing here can be held back behind an email address.
Each solving time is the interval between the frame that painted the sum and the timestamp the browser attached to your Enter key, kept in one array in this tab and used to compute twelve medians and two intervals. Nothing is written to storage, so a reload loses a finished run, and nothing about the sums or your answers is in the copied breakdown.