Skip to content
AbilityBench

Hiring assessment practice

Practice for the Watson-Glaser critical thinking appraisal

Twenty original items in twelve minutes, in the five sections the appraisal runs and in the order it runs them — free, no account, and the marks appear when the clock stops. What makes this exam awkward is not the difficulty of any one item but the five different answer scales it puts in front of you: a five-point inference judgment, then assumption made or not made, then follows or does not follow, then beyond reasonable doubt or not, then strong or weak. The report is built around the five-point scale, because that is where candidates lose most and a mark out of twenty cannot show it.

  • 100% free
  • No signup
  • 20 items, 12 minutes
  • All five sections
  • Original items, not the real bank

20 items in 12 m 00 s, in five sections that run in the appraisal’s published order. Every section asks a different question and offers a different answer scale, and switching between them is most of what makes the real thing hard.

  1. 1. Inference 5 items, 5 options

    Judge how far the facts in the passage carry each inference. Use nothing but the passage.

  2. 2. Recognition of assumptions 4 items, 2 options

    For each statement, decide whether the proposed assumption is one the speaker must be taking for granted.

  3. 3. Deduction 4 items, 2 options

    Take the premises as true and decide whether the conclusion follows from them necessarily.

  4. 4. Interpretation 3 items, 2 options

    Decide whether each conclusion follows beyond reasonable doubt from the paragraph.

  5. 5. Evaluation of arguments 4 items, 2 options

    Assume every argument is true. Judge it strong if it bears directly and importantly on the question, weak if it does not.

Two instructions that decide more marks than anything else

Take every passage, premise and argument as true, including the ones that are wrong about the world. Sections two to five are not asking whether a claim is correct; they are asking what follows from it, what it needs, or whether it bears on the question.

An argument can be true and weak at the same time. In the last section, strong means it bears directly and importantly on the question asked — not that it is factual, and not that it is persuasive.

The untimed pair covers the two scales that cost candidates most: the five-point inference scale and the strong-or-weak judgment. Both are worked through afterwards and neither reaches the mark. Answer with the number keys or the buttons.

How to practise the five sections

One clock, five answer scales, and two instructions that decide most of the marks.

  1. Learn the two scales that cost candidates most, untimed

    Two items run before the clock: one on the five-point inference scale and one strong-or-weak argument. Those are the two judgments people arrive without, and both are worked through afterwards. Everything else uses a two-way scale you will recognize on sight. Neither untimed item reaches the marks, and either can be skipped straight past.

  2. Work the sections in order, taking every premise as true

    Section one gives a plant's quarterly figures and five inferences to place on the scale. Sections two to five give statements, premises, a survey paragraph and a question under debate. Take all of them as true, including anything that is wrong about the world — none of these sections is asking whether a claim is correct, only what follows from it, what it needs, or whether it bears on the question asked. Answer with the number keys or the buttons; twelve minutes covers the whole set at an average of 36 seconds an item.

  3. Read the scale report, not the total

    The result gives a per-section table with two error directions — misses that accepted the affirmative option, misses that rejected it — and then measures the inference section in steps along its own scale. A one-step miss is a calibration problem at the True against Probably true boundary. A two-step miss has crossed Insufficient data, which means a passage was read as answering something it never touched. Those are different faults and the section total treats them identically.

Technical specifications

Items20 scored in five sections — inference 5, recognition of assumptions 4, deduction 4, interpretation 3, evaluation of arguments 4 — plus 2 untimed items outside the marks
Answer scalesFive options on inference (True, Probably true, Insufficient data, Probably false, False) and two on each of the other four sections. Five scales in one sitting is the structural feature of this appraisal, not a quirk of this page
Guessing floor8.5 of 20, since fifteen items are two-way. Worth knowing before reading a percentage off this kind of test: 12 of 20 is 60% and only 3.5 marks clear of blind guessing
Time12 minutes across the whole set — an average of 36 seconds an item, allocated as you like. Real sittings are timed per form and the limit is often set by the employer, so this budget is ours and is stated rather than borrowed
Section orderFixed, matching the appraisal's own sequence; only the items inside a section are shuffled. Interleaving the five scales would measure how fast you switch between them, which is a real cost and a different one
Items and their originAll 22 written for this page around two original stems — a plant's quarterly note and a district food-waste survey. No part of any published item bank appears here, and the appraisal's own items are withheld by its publisher precisely because circulated items stop measuring anything
Inference scale reportEvery miss measured as a distance in steps: one-step slips counted apart from misses that cross Insufficient data, plus the signed mean, which says whether the run sat toward True or toward False of where the passage put it
Reference figureNone. The publisher's norm tables are licensed to its customers, and several compare a candidate with applicants for one role at one employer rather than with a population, so no general average exists to print

Frequently asked questions

How many questions does the real appraisal have, and how long is it?

It depends on the form and often on the employer, which is why this page states its own count and clock rather than claiming to match. The appraisal is published in more than one length, and the sitting a candidate gets is configured by whoever bought it — some are strictly timed, some are generously timed, and some are administered untimed as a development exercise. What is stable across all of them is the structure: the same five sections in the same order with the same five response scales, and that is the part practice can transfer.

What separates insufficient data from probably false?

Insufficient data means the passage gives you nothing that bears on the inference in either direction; probably false means it gives you something, and that something runs against it. The test is not how likely the inference feels but whether the text moved the needle. An inference about pay, on a passage that never mentions pay, is insufficient data no matter how confidently you could guess it. An inference saying a change was the only cause of something, on a passage that names a second change in the same quarter, has been argued against — not settled, argued against — and that is the probably-false end.

How can an argument be true and still count as weak?

Because strength here means bearing on the question, and truth is not the same relation. The last section asks you to assume every argument is true and then judge whether it is important to the question and directly related to it. A true statement about a different requirement, a slogan that would support any proposal at all, an appeal to how long the present arrangement has lasted — each is unobjectionable and each leaves the question exactly where it was. Candidates who argue well for a living find this section the hardest, because they are used to judging arguments by whether they would work on an audience.

The deduction section marked a sensible conclusion as not following. Why?

Because deduction here means the conclusion cannot fail given the premises, and sensible is a much lower bar. A drifting reading really is often a dirty sensor; the premises just never said dirt was the only cause, so the inference is invalid however often it turns out right. The same applies in reverse: a conclusion can be absurd in the world and still follow, and then it follows. Holding those two apart under a clock, with a plausible answer in front of you, is the whole content of the section.

Is a score here comparable to a score from the real appraisal?

No, and treating it as one would be the single most misleading thing this page could let you do. The items are ours, written to the published format rather than taken from the instrument, so their difficulty is set by us and not calibrated against anything. The real appraisal reports against a norm group, and this page reports a raw count out of twenty with a guessing floor beside it. What carries over is familiarity: knowing that the inference scale has five points and where Insufficient data sits on it, and knowing to assume the premises, is worth marks on the day and does not depend on the items matching.

Which employers use it, and what is the pass mark?

Nobody outside a recruitment team can tell you either, and a page that names a pass mark has invented one. The appraisal is sold for roles where judgment under ambiguity is the job, and candidates most often report meeting it on graduate and lateral schemes in law, consulting and professional services — but whether a given firm uses it this year, and where they set the sift, is theirs. Thresholds are usually set on a percentile against a chosen comparison group, so the same raw score clears one competition and not another.

What happens if my tab goes to the background mid-item?

The item is struck out, the countdown stops, and the item rejoins the end of its section. That is deliberately the opposite of what an employer's delivery platform does: those keep counting and many flag the switch against your attempt, because they are policing an exam. This page is not policing anything, and a timed item you were not looking at is not a measurement — so it is thrown away and asked again rather than marked wrong. The result says how many items went that way.

What the appraisal is actually measuring, and why good arguers score badly

The five sections are not five topics; they are five kinds of restraint, and the appraisal is named for Goodwin Watson and Edward Glaser, whose framework it turns into items. Read the sections side by side and the same instruction runs through all of them: do not use what you know. Inference asks how far a passage carries a claim, and the middle of its scale is reserved for claims the passage does not touch. Recognition of assumptions asks what a speaker is committed to, not what would also be sensible to believe. Deduction asks what cannot fail, not what is likely, which is the one section that has a whole test of its own in the deductive-reasoning items. Interpretation asks what a set of figures forces, which is usually less than what they suggest. Evaluation of arguments asks whether a reason bears on the question, having already stipulated that it is true. None of the five rewards knowing more, and every one of them punishes filling a gap.

That is why the candidates who struggle are often the articulate ones. Somebody trained to build a case reads “the present timetable has worked for years” and hears an argument that would land in a meeting, so they mark it strong; the section wants to know whether it bears on the proposal, and it does not. Somebody used to giving advice reads a paragraph about a shorter working day and supplies the obvious fact that pay was protected, because a report that omitted it would be a bad report; the scale wants to know whether the text said so, and it did not. Both moves are professional competence in the wrong place. The practical consequence for preparation is that reading faster does not help here and reading more literally does — which is the opposite of the advice that works on the three-verdict verbal sets that graduate employers pair it with.

Two things go wrong on most free pages carrying this name. The first is collapsing the inference scale: a version that offers true, false and cannot say has removed the two calibration points the section exists to test, and a candidate who practises on it arrives having never had to choose between Probably false and Insufficient data. The second is worse and commoner — publishing what look like real items. The appraisal's item bank is withheld deliberately, and not only for copyright: an item whose answer is circulating measures recall rather than judgment, so a page that leaks items degrades the exam for everybody who sits it, including the person reading the page. Everything here is written for it, which is also why no percentile appears anywhere. For the same five judgments taken as a general ability, untimed and scored one section at a time, run the five sub-scale profile; for the wordless side of the same graduate sift, the figural contrast shows how much of a reasoning score is reading. The batteries employers pair with this one usually add matrix items and, for commercial roles, a forced-choice profile.

Not affiliated with, endorsed by, or connected to the owner of Watson-Glaser. Watson-Glaser is a trademark of its owner and is used here only to name the assessment this page prepares for. The questions on this page are our own: no part of the published test is reproduced, and a score here is not comparable to a score from the real instrument.

A reaction time here is the interval between the frame that painted the stimulus and the timestamp the browser attached to your key, both read from the same monotonic clock. What neither can see is the display pipeline behind it, so on a 60 Hz screen roughly 16 ms of every figure below is the machine rather than you. That is the timing floor: two numbers closer together than that are the same number, and this page reports no precision it cannot support.

Every figure on this page is in seconds or minutes, so that floor is four orders of magnitude below anything printed — but it is the reason the per-section medians are rounded rather than given to a decimal. What genuinely moves those medians is which section a passage sits in: the first item under a stem carries the reading of it, and on a three-item section that is a third of the column.

This is a measurement exercise, not a clinical assessment. It reports what you did on this page against a stated reference and nothing more — it cannot establish a result an employer would accept in place of their own sitting of the appraisal. Only a qualified professional, working with more than a browser, can make that judgment.

What happens to the twenty answers afterwards

Every number on this page is worked out by JavaScript running in the tab you are reading it in. Your answers, your reaction times and your score are never uploaded, logged or kept — which is also why the test carries on working after you disconnect from the network, and why nothing here can be held back behind an email address.

The twenty answers, the twelve-minute countdown and the section table are all built from one array inside this tab. No employer, recruiter or publisher can see that you took this, because nothing about it leaves the machine — and the copied summary carries section counts and step distances rather than which option you picked on any item.