Skip to content
AbilityBench

Free guide to the eight assessment families

What an employer can send you, and what each kind actually predicts

Eight families of assessment cover almost everything that arrives in a hiring email — a reasoning test, a work inventory, a scenario set, a work sample, an integrity questionnaire, a structured interview, a biodata form and an assessment center — and the table below orders them by how strongly each has been found to correlate with later job performance. It is free, there is nothing to sign into, and there is no exercise on this page: it is the map you read before you decide which of the drills is worth your evening. Both the 1998 estimates and the 2022 re-estimates are shown, because they disagree, and the disagreement is the most useful thing here.

  • 100% free
  • No signup
  • 8 families
  • 2 meta-analyses
  • No drill, no score

Eight families of hiring assessment, ordered by how strongly each one has been found to correlate with later job performance. Switch the ordering between the two meta-analyses and watch the list rearrange itself — that rearrangement is the single most useful thing on this page, because the version of this ranking most people have read is the older one.

Order the table by

Sackett, Zhang, Berry & Lievens (2022), Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range, Journal of Applied Psychology

Operational validity against rated job performance
Method20221998Mover² in 2022
0.420.51−0.0917.6%
0.380.35+0.0314.4%
0.330.54−0.2110.9%
0.310.51−0.209.6%
0.310.41−0.109.6%
0.290.37−0.088.4%
0.26not in the 1998 table6.8%
0.190.31−0.123.6%

Cognitive ability test

Speed of learning is the thing being bought. A short reasoning test predicts how quickly somebody gets to competence, which matters most where the job keeps changing.

The coefficient beside it is a correlation with training success and later job performance, pooled over the studies each meta-analysis could find. Between the two re-estimates it moved −0.20, and the reason is arithmetic rather than a change in the world: the same primary studies were re-corrected under a different assumption about who was in them.

See this format on this site — twelve-minute practice drill

Read the last column before you read the others

Square a correlation and you get the share of the differences between people it accounts for. The strongest method in the 2022 column squares to 17.6%, which means that even the best-evidenced instrument in the table leaves the large majority of the variation between two candidates unexplained. That is not an argument against testing — a hiring program choosing by that method still ends up materially better off than one choosing by chance, and across a hundred hires the difference is real. It is an argument against reading any single result as a verdict on one person, which is precisely what a candidate staring at their own number is tempted to do.

A second thing follows from it. Employers combine methods because the combination adds something none of the parts has on its own: a test of reasoning and a structured interview overlap much less than either overlaps with itself, so putting both in the process moves more than either would alone. If a process has sent you three different exercises, that is the reasoning behind it rather than indecision.

Where the situational judgement line comes from

Its absence from the 1998 column is not an omission on this page. The meta-analysis that established the family’s validity was published in 2001, and it is also the source usually quoted for the finding that how the instructions are worded — what you would do against what you should do — changes what the format ends up measuring.

McDaniel, Morgeson, Finnegan, Campion & Braverman (2001), Use of situational judgment tests to predict job performance: A clarification of the literature, Journal of Applied Psychology

Nothing on this page reads or writes anything about you: there is no input, no result and no token, and the ordering you chose is held in this tab until you leave it.

How to use this map before an assessment

Find the family, read what it claims, then go and sit the format once.

  1. Name the thing you were sent

    The invitation almost never says which family it belongs to, so work backwards from the shape. A clock and multiple-choice reasoning items is a cognitive ability test. Statements you rate as being like you, with no right answer visible, is a work inventory. A paragraph describing a colleague and four things you might do is a situational judgement test. A piece of the job itself, done under observation, is a work sample. The names publishers use are brand names; the families are what the research is about.

  2. Read the two coefficients, then the squared one

    Tap a method to open it. The 2022 figure is the current estimate, the 1998 figure is the one most articles still quote, and the last column squares the current one so you can see how much of the difference between two people it accounts for. Nothing in that column exceeds a fifth. Use it to calibrate how much weight the process is really putting on any one exercise, and how much is being decided elsewhere.

  3. Go and sit that format once, cold

    The links under each family go to the page on this site where that format is put in front of you. Doing one under a clock tells you something the description cannot: whether the unfamiliar part is the reasoning or the pace. That is the only claim practice makes here — a format you have seen before costs you less time on the day than a format you have not.

Technical specifications

Families coveredEight: structured interview, empirically keyed biodata, work sample, cognitive ability, integrity, assessment center, situational judgement, and conscientiousness from a personality inventory
Reference figuresTwo full sets of operational validity coefficients from independent meta-analyses, Sackett, Zhang, Berry & Lievens (2022), Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range, Journal of Applied Psychology, and Schmidt & Hunter (1998), The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings, Psychological Bulletin
Headline figureCognitive ability against rated job performance, r = 0.31 in the 2022 re-estimate against 0.51 in the 1998 review — the largest single fall in the table
The one that roseEmpirically keyed biodata, 0.35 in 1998 and 0.38 in 2022. It is the only family in the table whose estimate went up, and it did so because the range-restriction correction that inflated the others was doing comparatively little work on it
OrderingLive: the table re-sorts when you switch meta-analysis, and the top three change identity rather than merely shifting. Under the 1998 figures the work sample leads; under the 2022 figures it is fourth
Squared columnEach 2022 coefficient shown as r², from 17.6% at the top of the table to 3.6% at the bottom. Printed because a correlation read as a percentage is the single most common misreading of this literature
What is not hereNo percentile, no passing mark, no cutoff, and no figure for what any employer accepts. Those are set per employer per role, they are not published, and a number invented for that slot is the one thing on this site that could cost somebody a job
Outbound linksTwelve, all internal to this site: every family that can be practiced here is linked from the family that names it, and from the list at the foot of the panel

Frequently asked questions

Which of these will I be given?

Almost certainly a cognitive ability test, a work inventory, or both, because they are the cheapest to administer at volume and they arrive as a link rather than as a diary invitation. Work samples and assessment centers cost an employer real money and real hours, so they appear late in a process and usually only for roles where the cost is small next to the salary. A situational judgement test tends to arrive in graduate and high-volume schemes, where it is doing the work of a first sift. If you have been sent a link with a clock on it and no mention of a human being, the odds are heavily on the first family in that list.

Why did the 1998 numbers change so much in 2022?

Because of a statistical correction applied to the primary studies, not because anybody re-ran the studies. Validity research is nearly always done on people who have already been hired, which makes them less varied than the applicant pool the method is meant to sort — and a correlation computed on a narrowed group understates the one you would see across the whole pool. Everyone corrects for that. The 2022 argument is that the standard correction assumed a degree of narrowing that the data do not support, so the corrections were too large and the corrected coefficients too high. Re-doing them with a defensible assumption pulls the whole table down, and pulls some rows down much further than others.

Does a low coefficient mean the test is useless?

No, and this is where the arithmetic and the intuition come apart hardest. A correlation of 0.31 sounds like a rounding error and is in fact enough to change who gets hired in a measurable way when it is applied across hundreds of decisions, because the employer is not predicting one person — it is choosing the top slice of a queue. What a coefficient of that size will not support is a confident statement about the individual at the front of the queue, which is why an employer using these methods properly is running a program, not making an inference about you personally.

Are personality questionnaires worth taking seriously at all?

They are worth taking seriously as part of a process and not as a measurement of you. Conscientiousness sits at the bottom of this table because self-report about habitual behavior predicts rated performance weakly on its own, and because the ratings it is validated against are themselves noisy. What keeps these instruments in use is that they add something a reasoning test does not, and that they are cheap. The forced-choice format many of them use exists precisely because a rating scale invites everyone to be maximally agreeable and dependable.

Can an employer legally test me on this?

The relevant obligation in most jurisdictions is that a selection method be job-related and consistent with business necessity, which is why publishers sell validation studies alongside their tests. This page is not legal advice and does not pretend to be. What it can tell you is what a defensible process looks like from the outside: the same exercise given to everybody at the same stage, scored the same way, with the results used alongside other evidence. A test given to some candidates and not others, or one whose result is the whole decision, is unusual enough to be worth noticing.

Should I disclose a disability and ask for adjustments?

That decision is yours and this page will not push it either way, but two facts are worth knowing before you make it. Almost every commercial assessment publisher documents an adjusted administration — extra time is the common one — and the request goes to the employer rather than to the publisher. Second, a timed reasoning test given without an adjustment somebody needed does not produce a lower score for the reason the employer thinks it does; it produces a measurement of the barrier. That is a general problem with speeded formats and not a statement about you.

Where does this leave practice?

Practice buys familiarity with a format and nothing else, and familiarity is worth real minutes on a paper with a clock on it. What it does not buy is a different underlying ability, and any site telling you it does is selling something. The honest version of the claim is narrow: on a fifteen-minute paper where the average candidate never reaches the end, the seconds you save by not having to work out what an unfamiliar question type wants are seconds that go into answering questions instead.

What a validity coefficient is, and what a candidate should do with one

Selection research has one dependent variable and it is rarely stated out loud: a supervisor’s rating of an employee’s performance, collected some months after hiring. Every coefficient in the table above is a correlation between something measured before the hire and that rating afterwards. This matters because the rating is not a perfect record of how good somebody is at their job — it is one person’s judgment, and it carries its own error. A method that correlated perfectly with the true quality of an employee could not show a correlation of 1.0 with a noisy rating of it, so every number in the literature is bounded from above by the quality of the criterion rather than by the quality of the test. Meta-analysts correct for that too, which is another reason the corrections became the subject of the argument.

The 2022 re-estimate did more than lower the figures; it reordered them, and the new order tells a candidate something practical. The structured interview now heads the table, which means the most predictive part of most hiring processes is the part involving a human being asking everyone the same questions — not the test that arrived by email. The work sample fell from first place to fourth, which is the most surprising movement in the set given how intuitive the method is. And cognitive ability, the family this site has the most pages about, fell from 0.51 to 0.31. That last figure is the one to carry into an evening of practice: it is the size of the thing a reasoning test is measuring about the people it sorts, and it is smaller than the industry built around it implies. If a reasoning paper is what is in front of you, the mixed drill on this site will tell you which of the four question types costs you the most time, and the page that reads a raw count explains exactly where the arithmetic on such a count runs out.

One family behaves differently from the rest and deserves a paragraph of its own. A situational judgement test has no key until somebody builds one, and the way it is built is by asking experienced people in that job at that employer what they would do, then scoring candidates against the pattern of those answers. That makes the instrument valid in a way a generic version cannot be, and it also makes the format the only one in the table where preparation is genuinely about understanding the shape of the question rather than about speed — the scenario page takes that apart properly. Two of the largest publishers get their own pages for the same reason: what an adaptive battery does to the difficulty of your next question, and what a four-dimension talent model is trying to separate, are structural facts about those processes that no amount of practice on reasoning items will teach you.

Cognitive ability test, operational validity (r)

0.31Correlation between a short cognitive ability test and later supervisor-rated job performance, pooled across published criterion-related validity studies.

Sackett, Zhang, Berry & Lievens (2022), Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range, Journal of Applied Psychology

This figure describes how well the method sorts a pool of applicants when an employer uses it on hundreds of people. It is not a score, it cannot be applied to one person, and it says nothing about how any particular employer weighs a result. The earlier and much-quoted estimate of 0.51 came from the same body of primary studies corrected under a different assumption about how varied the people in them were.

This is a measurement exercise, not a clinical assessment. It reports what you did on this page against a stated reference and nothing more — it cannot establish a learning difficulty, an attention disorder, or anything else that would justify an adjusted administration. Only a qualified professional, working with more than a browser, can make that judgment.

What this page knows about your process

Every number on this page is worked out by JavaScript running in the tab you are reading it in. Your answers, your reaction times and your score are never uploaded, logged or kept — which is also why the test carries on working after you disconnect from the network, and why nothing here can be held back behind an email address.

This page in particular has nothing to keep: there is no exercise on it, so no answer, no timing and no result exists to store. The only state it holds is which of the two meta-analyses you asked it to sort by, and that lives in the tab until you close it.