Skip to content
AbilityBench

Vocabulary size

Vocabulary test with the guessing subtracted

Sixty strings appear one at a time and you say whether you know each one. Forty are English words and twenty were built for this page, which is how the count is kept honest: whatever share of the invented strings you claim is subtracted back out of the real words you claimed. Free, no account, and the corrected figure, the formula behind it and its error bars are all on screen at the end.

  • 100% free
  • No signup
  • 60 strings, 20 invented
  • Corrected for false alarms
  • Both rates shown

60 strings, one at a time. 40 of them are English words and 20 were invented for this page. Say yes only to the ones whose meaning you could give if somebody asked — the invented ones are how the page finds out whether you meant it.

What the two answers have to mean

Yes
you could define it, or use it in a sentence that shows the meaning. A vague feeling of having seen it is not a yes, and the invented words are built to produce exactly that feeling.
No
anything else, including a word you half-recognize. Saying no to a word you do know costs you a little; saying yes to one you do not costs you far more, because every invented word you claim is subtracted from your real words as well.

The warm-up shows two real words and two invented ones and tells you which was which afterwards. None of the four appears in the scored list, and no clock runs at any point in this test.

How to take the vocabulary test

One rule about what yes means, sixty strings, and one subtraction.

  1. Fix what your yes is going to mean

    Say yes only to a string whose meaning you could give if somebody asked — a definition, or a sentence that shows the sense. A feeling of having seen it somewhere is not a yes, and the twenty invented strings were assembled specifically to produce that feeling: they are pronounceable, they are spelled the way English spells things, and they carry ordinary endings. Four warm-up strings, two real and two invented, let you calibrate before anything counts.

  2. Work through sixty strings with J and F, or the two buttons

    One string at a time, no clock, no feedback. Nothing tells you whether a string was real until the run is over, because being told would let you recalibrate mid-test and the calibration is the thing being measured. The order is drawn from the run seed with no more than four of either kind in a row, so the sequence cannot be answered by rhythm.

  3. Read the two rates before the number

    The results panel shows your hit rate on the real words, your false-alarm rate on the invented ones, and the corrected proportion those two produce, in that order — the count of words comes after them because it depends on both. Under that sit the sensitivity figure, where you set your yes threshold, the invented strings you claimed, and the real words you passed on.

Technical specifications

Strings60 scored — 40 real words spread across five frequency tiers of eight, and 20 invented strings — plus 4 warm-up strings held outside the scored list. Every run shows the same 60 in a different order
How the invented strings were builtPronounceable, orthographically legal, and carrying ordinary English endings such as -ous, -ible and -ate, so none can be rejected on shape alone. A nonword that nobody could believe measures nothing, which is why they are not random letter strings
The correctionCorrected proportion = (hit rate − false-alarm rate) ÷ (1 − false-alarm rate). Anderson & Freebody (1983), Reading comprehension and the assessment and acquisition of word knowledge. It asks what share of the words you did not know you claimed anyway, and removes that share from the words you did claim
The reference figure42,000 lemmas, the average for an American twenty-year-old — Brysbaert, Stevens, Mandera & Keuleers (2016), How many words do we know? Practical estimates of vocabulary size dependent on word definition, the degree of language input and the participant's age, Frontiers in Psychology. The same work puts a sixty-year-old at 48,000, which is why the page reports a share of a stated reference rather than a bare word count
What a lemma isA base form with its inflections counted once, so walk, walks, walked and walking are one lemma and not four. Estimates quoted elsewhere run two to three times higher because they count word forms, and a comparison between the two kinds of figure is meaningless
The error barsA 95% Wilson interval on the 40 real words, carried through the same correction and then scaled. It covers the luck of which 40 words you happened to be asked and nothing else — not the choice of reference figure, which is the larger uncertainty and has no interval anybody can compute
Sensitivity and biasThe four cells are also reported as d-prime and a criterion, computed with the log-linear correction so that a flawless run gives a large finite number rather than infinity. A negative criterion is a lean toward saying yes and is exactly what the subtraction removes
What leaves the pageNothing unless you press copy. The copied summary carries both rates, the corrected figure and the run seed, and never a list of which strings you personally claimed

Frequently asked questions

Why is a third of the list made up?

Because without them the test measures how willing you are to click yes. Nobody can check a claim to know a word from the claim alone, so a plain checklist rewards confidence and punishes caution, and two people with identical vocabularies can come out 30% apart on nothing but temperament. The invented strings give the page a baseline: the rate at which you claim things you cannot possibly know is measurable, and once it is measured it can be taken out of the rest of your answers. That single design decision is the difference between a vocabulary estimate and a personality quiz.

I said yes to a word I only half-knew. Have I broken the test?

No, you have fed it. Half-knowing is the normal state of most of anybody's vocabulary, and the correction is built on the assumption that half-known real words and plausible invented ones get claimed at similar rates. If you are generous with your yes answers, you will be generous on the invented strings too, and the subtraction removes both together — which is why the results panel shows your criterion, the number that says which way you leaned. Someone who leaned toward yes and someone who leaned toward no can land on the same corrected figure, and that is the correction working.

Why can the estimate never come out above the reference figure?

Because the estimate is a share of that reference, not an independent count. The forty words are treated as a sample of the 42,000 lemmas an average American twenty-year-old holds, so knowing all forty says you hold that reference vocabulary and the arithmetic stops there — the page cannot see past its own hardest word. Sites reporting 68,000 or 80,000 have extrapolated beyond their item list, usually by assuming that the rate at which you knew their rare words continues into rarer words they never asked about. That assumption is not testable from the answers they collected.

Is walk, walked and walking one word or three?

One, on the count this page uses, and three on the count most other sites use. A lemma is a base form together with its regular inflections, which is the unit the reference figure was collected in, so the two have to match or the comparison is nonsense. This is the single commonest reason two vocabulary tests disagree by a factor of three about the same person: not that one is stricter, but that they are counting different things and neither says which. Compound and derived forms are a further complication — whether unhappiness is its own lemma or a form of happy is a decision, and different published counts decide it differently.

Another site gave me a much bigger number. Which one is right?

Ask both what they multiplied and by what. Every word-count estimate is a proportion times a reference size, and the only thing that distinguishes a defensible one from a decorative one is whether it says which proportion and which reference — this page prints the formula, the reference, the source of the reference and the interval, so the arithmetic can be checked line by line. A number offered without those cannot be wrong in any useful sense, because there is nothing in it to disagree with. It is also worth checking whether the other test had any nonwords in it at all; if not, its number is your willingness to click yes, dressed up.

Does vocabulary really keep growing with age?

Yes, and it is one of the few measures on this site that moves that way. The same work behind the reference figure puts a twenty-year-old at 42,000 lemmas and a sixty-year-old at 48,000 — roughly a word and a half a day over four decades. That matters for reading your own result: a page that compares everybody to a single average is comparing a nineteen-year-old and a retiree to the same thing, and one of them is being told something false. This test does not ask your age and therefore does not adjust for it, which is a limitation stated rather than a feature.

Can I use this as proof of my English level?

No, and the reason is worth knowing if you are studying for a placement. CEFR levels are descriptors of what a learner can do, not score bands. Every mapping from a test score to a level is that test's own calibration decision, and different published mappings disagree about where B2 begins. What a size estimate can tell you is how much of the lexicon you are carrying, which is a real input to reading comfort and a poor substitute for a placement decision. Nothing here should be sent to an admissions office, and a number from this page will not be recognized by one.

The yes/no checklist and the hole in the middle of it

Showing somebody a list of words and asking which ones they know is the cheapest vocabulary measurement in existence, and for most of a century it was also one of the least trusted, for an obvious reason: the answer is unverifiable. There is no way to check a yes. The fix that made the format usable is not a better question but a different kind of item — strings that cannot be known, mixed in where they cannot be spotted, so that the rate of unearned yes answers becomes visible and can be removed. The arithmetic this page uses is Anderson & Freebody (1983), Reading comprehension and the assessment and acquisition of word knowledge: subtract the false-alarm rate from the hit rate and divide by what is left, which treats an unknown word as a coin the reader flips at their own personal rate. It is not the only published treatment, and the alternatives disagree on the same answers — averaging the two percentages instead of dividing, as one widely used lexical test for advanced learners does, is gentler on a reader who claims a lot and harsher on one who claims nothing.

The step after the correction is where most word-count tests quietly stop being measurements. A corrected proportion is a number about the list you were shown; turning it into a count of words requires multiplying by a vocabulary size that came from somewhere else, and that reference is then part of the claim rather than a detail. Brysbaert, Stevens, Mandera & Keuleers (2016), How many words do we know? Practical estimates of vocabulary size dependent on word definition, the degree of language input and the participant's age, Frontiers in Psychology is where the 42,000-lemma figure on this page comes from, collected from a very large online sample answering exactly this kind of checklist, and it carries its own conditions: it counts lemmas rather than word forms, it is American English, and it moves with age — 48,000 lemmas by sixty. Multiplying by it imports all three. The visible consequence is the ceiling: because the forty words are treated as a sample of that reference vocabulary, a perfect run reports the reference and not more, and the page says so instead of extrapolating past the last word it asked about.

What a size estimate is good for is narrower than it looks, and worth being clear about before reading yours. It predicts how much unfamiliar vocabulary a text will throw at you, which is why it belongs next to the reading level test rather than next to an intelligence measure. It says nothing about whether you can produce those words, spell them, or use them in the right register — the written half is the spelling test, and the meaning-under-pressure half, where four options force you to commit to a sense, is the synonym test, which also hands back the definition of everything you miss. If your reason for arriving was a placement rather than curiosity, the English level test covers grammar and reading alongside words, and the reading age test reports the band a school would use. Recognizing a word is also only half of holding it, and the half no checklist can reach is what a word drags up unprompted, which is the entire business of the word association test — a page built with no answer key at all.

A reaction time here is the interval between the frame that painted the stimulus and the timestamp the browser attached to your key, both read from the same monotonic clock. What neither can see is the display pipeline behind it, so on a 60 Hz screen roughly 16 ms of every figure below is the machine rather than you. That is the timing floor: two numbers closer together than that are the same number, and this page reports no precision it cannot support.

No result on this page is a time and none of the sixty strings is raced. The floor earns its place here through one validity rule: a keypress landing within 150 ms of a string being painted is thrown out, because at sixty items in a row the commonest cause of a figure that small is a finger still travelling from the previous answer rather than a judgment about the string in front of it.

This is a measurement exercise, not a clinical assessment. It reports what you did on this page against a stated reference and nothing more — it cannot establish a language impairment, a learning disability or a reading delay. Only a qualified professional, working with more than a browser, can make that judgment.

Where the sixty answers are counted

Every number on this page is worked out by JavaScript running in the tab you are reading it in. Your answers, your reaction times and your score are never uploaded, logged or kept — which is also why the test carries on working after you disconnect from the network, and why nothing here can be held back behind an email address.

The word lists, the invented strings and the correction are all in the JavaScript this page loaded, so the two rates are computed beside you and no answer is transmitted anywhere. One consequence is worth knowing: because nothing is stored, a second run cannot be compared with your first by the page, and because you now know which strings are invented, it would not be an independent measurement if it could.