Skip to content
AbilityBench

Comprehension

Reading comprehension test scored in four separate kinds

Two passages, twelve questions, and four sub-scores rather than one total — free, no account, and nothing kept back at the end. Three questions ask what the text states, three what it implies, three what a word means in the sentence it sits in, and three what the passage as a whole was for; the panel reports each kind on its own with the interval three items leaves around it. The passages stay in front of you the entire time, because a question you can only answer from memory is measuring memory.

  • 100% free
  • No signup
  • 12 questions, 4 kinds
  • Passages stay on screen
  • Wrong answers explained

Two passages — a piece of history and a short argument — and 12 questions between them, 3 of each of four kinds. Both passages stay on screen the whole time. Nothing is timed against you, and the result is reported kind by kind rather than as one number out of 12.

The four kinds of question

Stated in the text
The answer is a phrase you could underline. Missing these usually means the eye skipped rather than that the reader failed to understand.
Inferred from the text
Nothing in the text says it; the text makes it the only sensible reading. This is the kind that separates readers, and the one most browser tests leave out because it is the hardest to write a defensible key for.
Word in context
A word whose everyday sense is not the one the sentence needs. The passage supplies enough to fix the meaning, so knowing the word beforehand is not what is being asked.
Point of the whole
What the whole passage was for, as opposed to what it happened to mention. The wrong answers here are all true statements about the text, which is what makes the item hard.

Three items per kind is few, and the panel at the end prints the interval around each sub-score rather than pretending otherwise. Three out of three and one out of three are far enough apart to mean something; two and three are not, and the page will say so.

How to take the comprehension test

Two passages, four kinds of question, and a reason attached to every answer you miss.

  1. Read the report, then answer six questions about it

    The first passage is 236 words of history about the mark on the side of a cargo ship. It stays in a scrolling panel above the question for as long as you want it, so going back to check a clause is expected rather than penalized. The six questions are shuffled by a seed, and the kind each one belongs to is printed above it — you can see whether you are being asked what the text said or what it meant.

  2. Read the argument, then answer six more

    The second passage is 259 words making a case, which changes what the questions can ask. A report can be checked against itself; an argument has a position and evidence offered for it, and the items are largely about telling those apart. Nothing is timed, nothing is cut off, and there is no penalty for a slow answer — the run records how long each item took, but only to show you which kind you sat on.

  3. Read the four sub-scores, and the reasons for the misses

    The result gives correct out of twelve with a Wilson interval, the exact probability of that total by guessing at four options, and then the same figures for each kind separately. Every item you got wrong is listed with the answer and a paragraph explaining why that answer and not yours — which is the part worth reading, since a sub-score of one out of three tells you where to look and the explanation tells you what to look at.

Technical specifications

Items12 across 2 passages: 3 stated-in-the-text, 3 inference, 3 vocabulary-in-context, 3 main-idea. Four options each, one defensible key, item order shuffled from a seed the result carries
Passages236 words of narrative history and 259 words of argument, both written for this page. Word counts are taken from the strings that are rendered, so the figure on screen cannot drift from the text above it
Passage availabilityPermanently on screen in a scrolling panel while the questions run. This is the opposite of the choice made on the reading speed test next door, and it is the difference between measuring comprehension and measuring what survived a single pass
ScoringCorrect out of scoreable, with a 95% Wilson interval that does not collapse to zero width at 12 out of 12, plus the exact one-sided binomial probability of the total against a 25% guessing floor
Sub-score resolution3 items per kind. A one-item gap between two kinds is inside the noise and the page says so; a two- or three-item gap is the pattern the test is arranged to expose
TimingNothing is speeded and no response window closes. Per-item intervals are recorded and reported as medians in seconds, per kind, over correct items only — a secondary readout, not a score
What is not reportedNo grade, no reading age, no percentile. There is no published distribution for twelve browser-administered items written for one page, and a rank derived from anything else would be arithmetic performed on the wrong sample
What leaves the pageNothing. The copy button puts the four sub-scores and a seed token on your own clipboard; the passages, your answers and the intervals stay in the tab

Frequently asked questions

Isn't leaving the passage on screen just letting people look up the answers?

Looking things up in the passage is reading, and the alternative measures something else. A comprehension question is not a memory question: a reader who has followed an argument and cannot quote its third sentence has still followed the argument. There is a mechanical reason too — in ordinary skilled reading about 10-15% of eye movements go backwards over material already covered, per Rayner (1998), Eye movements in reading and information processing: 20 years of research, Psychological Bulletin, which means re-reading is part of how comprehension is built rather than a way around it. What stops the items being lookups is how they are written: the inference and main-idea items have no sentence you can point at, and the vocabulary items give you a word whose everyday sense is the wrong one.

How do you know the answer to an inference question is the right one?

By making the wrong options true. The failure mode of a badly written inference item is that two of its options are defensible, so the key is really the writer's preference; the way out is to build every distractor from something the passage actually says, and then make only one of them the thing the passage is organized to establish. Each item here is published with the reasoning attached, in the panel at the end, so the key is checkable rather than asserted — if the argument for an answer does not convince you, you can see exactly where it went wrong instead of concluding you failed.

Three questions per kind is not many. Why so few?

Because the alternative is a test long enough that most people abandon it, and a sub-score from an abandoned test is worse than a coarse one. Three items resolves a difference of two or three items and nothing finer, which is why the result panel prints an interval next to every sub-score instead of a bare fraction: at three items a score of two out of three has a 95% interval running from roughly a fifth to nearly all of the scale. Read the sub-scores as a shape rather than as four numbers, and rerun the test if the shape surprises you.

Why is there no reading age or grade level attached to my score?

Because a grade level is a property of a text and not of a person, and a reading age is a score on a standardized school assessment rather than something twelve questions on a web page can produce. Those two get answered properly elsewhere on this site rather than approximated here. What this page can honestly report is which kind of question your errors landed in, and that needs no scale behind it.

Does it count against me if I take a long time on a question?

No — nothing on this page is timed against you, and no response window ever closes. Each item's duration is recorded because the four kinds usually separate on it, and inference items are typically the ones people sit longest on, which is worth seeing beside a low inference sub-score: slow and right is a different situation from fast and wrong. Speed is measured properly on a different page, where it is the point rather than a side effect, and where the passage is taken away to make the check mean something.

I scored twelve out of twelve. What does that tell me?

It tells you the test ran out before you did, which is a real result but a bounded one. A ceiling score means only that the hardest item here was not hard enough to find your limit, and this page has no adaptive tail to go and look for it — the difficulty of the passages is fixed. If you want to know where your ceiling actually is, climb a ladder of progressively harder text instead. The one thing a perfect score does not mean is that you read at any particular rate: comprehension and speed come apart above 300-500 wpm, and this test has nothing to say about which side of that you are on.

Are these like the reading questions on an admissions or language exam?

The item formats are the same family — a passage, four options, a defensible key — and the four kinds used here appear under various names in most published comprehension sections. What is not the same is everything that makes a published exam a measurement: a bank of hundreds of items trialed on thousands of candidates, a difficulty estimate for each one, and a score scale built from those trials. Nothing on this page is drawn from any published item bank, and a score here converts to no exam's scale. Treat it as a shape of strengths, not as a prediction.

What a comprehension question can actually establish

There is no agreed taxonomy of comprehension items — different testing traditions cut them into three, four or six categories and disagree at the joins — but almost all of them separate the same first distinction, and it is the one that carries the information. Some questions have an answer you could underline: the passage states it, and getting it wrong usually means the eye skipped rather than that anything was misunderstood. Other questions have no underlinable answer at all. The text supports exactly one reading, and the reader has to hold two or three sentences together to see which. Adding those two into a single score out of twelve throws away the only distinction the test was in a position to make, which is why this page never prints that total on its own. The other two kinds here — a word whose sentence forces a sense different from its everyday one, and a question about what the whole passage was for — fail in their own characteristic ways: the first catches readers who fill in a familiar meaning without checking it against the sentence, and the second catches readers who remember what a text mentioned but not what it was arguing.

The second thing worth knowing is what the distractors are doing. In a weak item the three wrong options are wrong because they are false, which turns the question into a fact check; in a strong one they are wrong despite being true, so the reader has to decide which true statement the question is actually asking about. Every main-idea item on this page is built that way, and it is the reason a main-idea score can sit below a detail score in somebody who reads perfectly well: they are answering a slightly different question from the one on screen. Where an answer is missed, the panel at the end sets out why the key is the key, because an item whose reasoning cannot be written down in three sentences is an item that should not have been asked.

What this test deliberately does not do is convert any of it into a level. The reading level test takes that job on properly by raising the difficulty of the text until accuracy gives way, and the reading age page is written for the parent or teacher who needs to judge a book rather than a reader. Errors that cluster in the vocabulary items usually point somewhere else again — to how much of the language is available without effort — and that is measured directly by the synonym test and the sentence completion test, which puts a missing word back into a sentence that constrains it. If the inference items were the ones that went wrong, the interesting neighbor is the verbal composite, where the same operation appears without a passage wrapped around it.

This is a measurement exercise, not a clinical assessment. It reports what you did on this page against a stated reference and nothing more — it cannot establish a reading difficulty, a language disorder or an attention problem. Only a qualified professional, working with more than a browser, can make that judgment.

Where your twelve answers are scored

Every number on this page is worked out by JavaScript running in the tab you are reading it in. Your answers, your reaction times and your score are never uploaded, logged or kept — which is also why the test carries on working after you disconnect from the network, and why nothing here can be held back behind an email address.

The four sub-scores, the intervals around them and the list of items you missed are all computed from an array that never leaves this tab, and the passages themselves are part of the page rather than something fetched while you read. Closing the tab discards the run, which is why the copy button exists and why there is nothing to delete afterwards.