Skip to content
AbilityBench

Three levels, one keypress

Language processing test that names the level, not a score

Everything this test shows you is a claim, and you press one key if it is true and another if it is not — 42 claims in about four minutes, free, with no account and nothing withheld at the end. What changes between the three sections is the level of language the claim lives at: a word and the category above it, a sentence with the word not in it checked against a picture, and a sentence with a whole clause buried inside it. Because all three use the same two keys, the three costs it reports can be compared with one another, and the page finishes by pointing at whichever of eight other tests measures your slowest level properly.

  • 100% free
  • No signup
  • 42 claims, ~4 minutes
  • 3 levels compared
  • Points you to 8 tests

42 scored trials in three sections, about four minutes. Everything on screen is a claim, and you press J if it is true and F if it is not. What changes between sections is the level of language the claim lives at, and the run reports what each level cost you.

Words and categories

18 claims, up to 4 s each

Sentences against a picture

12 claims, up to 8 s each

Sentences that embed a clause

12 claims, up to 12 s each

What this is for, stated up front

Six trials a cell is enough to separate three levels inside one person often enough to be useful and nowhere near enough to place that person against anyone. The output is an ordering of three of your own numbers and a pointer to the page that measures each one properly — a short run that tells you where to spend a longer one. It is not a language assessment and nothing here produces a composite score.

The practice covers two claims of each of the three kinds, has no time limit on any of them, and says after each one whether the key was right. None of its six sentences appears again in the scored run.

How to take the language processing test

One question, three levels, and a direction at the end rather than a verdict.

  1. Learn the two keys on six practice claims

    J means the claim is true and F means it is false, and those two keys stay the same through all three sections so nothing has to be relearned mid-run. The practice block covers two claims of each kind, has no time limit on any of them, and says after each one whether you were right. Its six sentences appear nowhere in the scored run, so nothing you have already worked out comes back as a familiar item.

  2. Work through the three sections in order, resting between them

    Section one is eighteen short sentences about what belongs to what, with up to four seconds each. Section two is twelve sentences checked against two stacked shapes, with up to eight seconds each, and half of them contain the word not. Section three is twelve sentences with a clause inside them plus a plain claim underneath, with up to twelve seconds each. A screen between sections tells you what is coming and nothing is timed while it is up.

  3. Read the three costs, then follow the pointer

    Each section yields one difference in milliseconds: two category steps minus one, a negated sentence minus a plain one, and an object-relative clause minus a subject-relative one. Each comes with a 95% interval, and an interval containing zero is labeled as a section that did not separate. The page then names your largest cost and links to the test that measures that level with enough trials to mean something — which is what the four minutes were for.

Technical specifications

Trials42 scored — 18 category sentences, 12 sentence-and-picture items, 12 embedded-clause items — plus 6 optional practice claims. Six trials per scored cell, which is stated in the results panel as the reason nothing here is a score
One response channelEvery claim in all three sections is answered with the same two home-row keys, read by physical position so they are the same two keys on QWERTY, AZERTY and Dvorak. That is what makes the three costs comparable: the constant price of reading the mapping and moving a finger is identical in all six cells and cancels out of every subtraction
Section one, word to categorySix sentences one category step up, six two steps up, and six false fillers so that answering true to everything scores 67% rather than 100%. Twelve different subject nouns, not the same six at two distances, because reading one noun twice in a four-minute block primes the second reading
Section two, negationThree shape pairs drawn as outlines with no color coding, one arrangement each way. The affirmative and negative sentences of a pair differ in exactly one word, and each condition carries three true and three false answers so the key gives nothing away about whether the sentence was negated
Section three, clause embeddingTwelve different noun pairs, six subject-relative and six object-relative, each with a bare transitive claim underneath in the identical form. No noun pair is read twice, so the only difference between the conditions is where the buried clause sits
Response windows4 seconds for a short category sentence, 8 for a sentence plus a picture, 12 for a sentence with a clause inside it — set by how long the material takes to read rather than by one number applied to all of it. Fixation is 600 ms and the gap before each claim is drawn between 400 and 800 ms
What voids a trialA response inside 150 ms of the sentence painting, a response before it paints at all, and any trial during which this tab went to the background. On sentence material that last rule matters more than elsewhere, because looking away for two seconds is longer than the reading itself
What is not herePicture naming, which is the fourth section a full battery would run. Naming latency is measured from the onset of the voice, this site asks the browser for no permission of any kind including the microphone, and a keyed choice among four names is not naming — so the section is absent rather than faked

Frequently asked questions

Is this the Language Processing Test that a speech therapist uses?

No, and the difference is worth being blunt about. There is a published clinical instrument with a very similar name, sold as a kit, administered one-to-one by a speech-language pathologist, aimed at school-age children, and scored against norm tables that are part of what the kit costs. This page is none of those things: it is a four-minute reading-time comparison for adults, its items are our own, it produces no standard score, and it is not affiliated with the publisher of that instrument in any way. If a school or a clinician has recommended the real assessment, this page is not a substitute for it and cannot stand in for a referral.

Why does the word not make a sentence harder to check?

Because it has to be applied after the rest of the sentence has been understood. To decide whether the star is not above the plus, you first work out whether the star is above the plus and then flip the answer, and that extra step takes measurable time — it is one of the oldest and most reproducible findings in sentence processing. The section is built to isolate it: the affirmative and the negative version of each item differ in exactly one word, the picture is identical between them in half the pairs, and both conditions carry the same number of true and false answers, so the cost that comes out is the negation and not the difficulty of the picture.

What makes an object-relative sentence harder than a subject-relative one?

The distance between a noun and the verb that needs it. In the banker that the lawyer advised left the meeting, you meet banker, then have to hold it unfinished while a whole clause goes past, and only then find out what the banker did. The subject-relative version, the banker that advised the lawyer left the meeting, lets each noun meet its verb immediately. The gap is one of the largest effects in the reading-time literature and is why this construction is used to study working memory in comprehension. It is also why the response window in this section is three times the one in section one: the sentences are genuinely slower to read, and cutting them off would measure the window rather than the reader.

Six trials per cell is not many. Why so few?

Because the page is honest about what it is for. A run long enough to give a stable estimate of any one of these three effects would take twenty to forty minutes and produce one number; this run takes four and produces an ordering of three, which is the thing you actually need to decide what to do next. The consequence is stated everywhere it matters: each difference is printed with a 95% interval, an interval that contains zero is labeled as a section that did not separate, and the page says in the results panel that running it again can change the order. What it never does is round that instability off and present the ordering as a profile of you.

One of my costs came out negative. Is that a bug?

No, it is what six trials a cell looks like. A negative cost means the harder condition happened to be faster for you in this run, which occurs regularly when the true effect is smaller than the trial-to-trial noise — and the interval printed beside it will straddle zero, which is the page telling you exactly that. The one case worth a second look is a large negative cost with poor accuracy in the same section: if you were right on half of the twelve, the reaction time is over the half you were right on, and both numbers are describing a section you guessed rather than read.

Does the category section still hold up? I have read that it was overturned.

Half of it. The observation is solid — sentences that cross two category steps really do take longer to confirm than sentences that cross one, and you can watch it happen in your own eighteen trials. What did not survive is the explanation, which was that the categories are stored in a hierarchy and the extra time is the extra step being walked. Smith, Shoben & Rips (1974), Structure and process in semantic memory: A featural model for semantic decisions, Psychological Review demolished that by finding cases where the two-step sentence is faster: a chicken is an animal beats a chicken is a bird, because a chicken is a poor example of a bird. So the section measures a real latency difference whose cause is contested, and the page prints both citations rather than only the flattering one.

What does a screen reader do with the shapes in section two?

It reads a description — a star above a plus — which is correct and also changes what the section is measuring. Checking a sentence against a picture and checking a sentence against a second sentence are not the same operation, and the negation cost that comes out of the second one is a different quantity. Nothing about the section is color-coded, so a color-blind visitor is measuring exactly what everybody else is; the shapes are outlines told apart by form alone. If you are using a screen reader, the two sections either side of it are the ones to trust here, and the tests linked at the end of the run are all text.

Why a short battery is a router and not an assessment

“Language processing” is not one thing, and that is the whole reason this page exists in the shape it does. Getting from ink to meaning involves recognizing a word, retrieving what it means, working out what the sentence built from those words claims, and holding parts of that sentence in mind while the rest arrives — and those steps come apart in people. A reader can be quick at single words and slow at embedded clauses, or the reverse, and a single composite score adds those apart-coming things together and hands back a number that describes nobody. So this run does not compute a composite. It measures one difference at each of three levels, on the same two keys, and reports the three separately with an interval on each. Because the response channel is identical across sections, the constant cost of deciding and pressing is the same in every cell and drops out of all three subtractions — which is the only reason three numbers taken from three different kinds of material can be lined up next to each other at all.

Each section is a paradigm with a real history and a real departure from it. Section one is category verification, from Collins & Quillian (1969), Retrieval time from semantic memory, Journal of Verbal Learning and Verbal Behavior — where the latency difference was first reported and explained by a hierarchy that later work took apart. Section two is sentence-picture verification, from Clark & Chase (1972), On the process of comparing sentences against pictures, Cognitive Psychology, which used a star and a plus exactly as this page does; the departure is that the original presented the sentence first and the picture afterwards, while this version paints both together, so the cost measured here includes the comparison rather than only the encoding. Section three is the subject-relative against object-relative contrast that King & Just (1991), Individual differences in syntactic processing: The role of working memory, Journal of Memory and Language used to argue that comprehension differences are working-memory differences; the departure there is bigger, since the published work read the sentences word by word and this page shows the whole sentence with the probe under it, which folds reading and verification into one interval. In all three cases the departure is the reason the absolute numbers are not comparable to a published figure, and the difference between two of your own conditions is.

The fourth section a full battery would have is picture naming, timed from the moment the voice starts. It is absent, and the reason is a design commitment rather than an oversight: this site asks the browser for no permission of any kind, which includes the microphone, so there is no way to detect a voice onset here. The available substitute is to show a picture and have the visitor pick its name from four options — but that measures a four-way choice, not naming, and calling it naming latency would be the kind of quiet substitution that makes a whole site untrustworthy. What is measured instead is what the router is for: the table below is on this page before you press Start, because the useful output of four minutes is knowing where to spend twenty.

Each level, and the page that measures it with enough trials

  • Single-word recognition: Lexical decision task80 trials over four frequency bands, which turns a single latency into a curve.
  • Word retrieval under a clock: Verbal fluency testSixty seconds of production, scored for the runs of related words and the jumps between them.
  • How many words you know at all: Vocabulary testAn estimate of size with a correction for guessing, which a speeded task cannot give.
  • Sense relations between words: Synonym testAdaptive items that converge on your level instead of marching through a fixed list.
  • Rate through continuous text: Reading speed testWords per minute with a comprehension check, because a rate without one is a scrolling speed.
  • Evaluating what a passage argues: Critical thinking testInference, assumption and argument evaluation scored as separate sub-scales.
  • Producing a word that fits a frame: Sentence completion testGrammar and register judged in context, which is the reverse of judging a claim already written for you.
  • The plain speed underneath all three: Digit symbol substitution testNinety seconds of coding against published age bands, for when every level here came out slow together.

Two more sit one step further out, and both use a timed judgment rather than language: the Wisconsin card sorting test for the flexibility that switching between rules needs, and the balloon analogue risk task for the decision-making side, which is the one component of this list that has nothing to do with reading at all.

A reaction time here is the interval between the frame that painted the stimulus and the timestamp the browser attached to your key, both read from the same monotonic clock. What neither can see is the display pipeline behind it, so on a 60 Hz screen roughly 16 ms of every figure below is the machine rather than you. That is the timing floor: two numbers closer together than that are the same number, and this page reports no precision it cannot support.

This is a measurement exercise, not a clinical assessment. It reports what you did on this page against a stated reference and nothing more — it cannot establish a language disorder, aphasia, or a developmental language disorder. Only a qualified professional, working with more than a browser, can make that judgment.

Not affiliated with, endorsed by, or connected to the owner of Language Processing Test 3. Language Processing Test 3 is a trademark of its owner and is used here only to name the instrument this page is telling you it is not. The questions on this page are our own: no part of the published test is reproduced, and a score here is not comparable to a score from the real instrument. That instrument is a kit-based battery administered one-to-one by a speech-language pathologist to school-age children; this page is a four-minute adult reading-time comparison, and the resemblance is a phrase in English and nothing else.

Where your 42 judgments are worked out

Every number on this page is worked out by JavaScript running in the tab you are reading it in. Your answers, your reaction times and your score are never uploaded, logged or kept — which is also why the test carries on working after you disconnect from the network, and why nothing here can be held back behind an email address.

The sentences are compiled into the page, so a finished run has made no network request of any kind and works offline once loaded. The timestamps go into a JavaScript array in this tab, are used to form three subtractions, and are gone when the tab closes — which is why a reload loses a finished run and why the copy button exists.