Skip to content
AbilityBench

Free scenario drill scored against an expert key

Situational judgement test with the reasoning shown instead of a score

Six workplace scenarios, four possible responses each, and your job is to rank them from most to least effective — free, untimed, with no account and nothing held back at the end. What you get back is not a mark: after each scenario the page shows where its own reasoning puts each response and why, plus the competency an employer’s key would file that option under. There is no percentage anywhere on this page because a situational judgement key is built from the answers of experienced people in one specific job at one specific employer, and no page outside that employer has one.

  • 100% free
  • No signup
  • 6 scenarios
  • 24 responses explained
  • Untimed

Six workplace situations, four possible responses each, and your job is to put the four in order from most to least effective. Nothing is timed and nothing is marked. When you have ordered all four, the page shows you where this site’s own reasoning puts them and why, one option at a time.

Read this before you start, because it changes how you should read the answers

A real situational judgement test is scored against a key, and the key is built by putting the same scenarios to experienced people doing that specific job at that specific employer and recording how they rank the options. The candidate is then scored on how closely their ranking matches the pattern those people produced. That is the whole mechanism, and it is why the format predicts anything at all.

No page on the open internet can have that. There is no panel behind these six scenarios, no employer and no job — so there is no key, and a site that hands you a percentage on a generic scenario set has invented the standard it marked you against. What you get here instead is an argument you can disagree with: our ordering, written down in advance, with the reason for each placement stated so you can see where the reasoning would change if the job did.

Number keys 1 to 4 pick from the options still unplaced, top to bottom; Backspace takes back your last placement. Everything works by tapping as well.

How to work through the six scenarios

Rank four, read the reasoning, then argue with it.

  1. Rank all four responses before anything is revealed

    Pick the response you think is most effective, then the next, until all four are placed. Number keys 1 to 4 choose from whatever is still unplaced, top to bottom, and Backspace takes back your last pick; every option is also a button, so a phone works the same way. The four options are shuffled on every run, so position on screen carries no information about where anything belongs.

  2. Read all four explanations, not just the ones you got wrong

    The reveal orders the options by where this site's reasoning puts them and marks your own placement beside each. The paragraph under each option is the argument for that placement, and the line under that names the kind of response it is — deal with it yourself, go to the person, escalate, or hold — and the competency a real key would file it under. The interesting disagreements are usually between the second and third place, where the argument is genuinely close.

  3. Look at your own pattern at the end, and distrust it

    The final panel counts your twenty-four placements by kind of response: how often you put escalation first, how often you put it last. That is a description of six decisions and nothing more — six scenarios cannot establish anything stable about a person, and the same six answers would be read differently by an employer hiring for a regulated role. Use it to notice a habit, not to conclude one.

Technical specifications

Scenarios6, written for this page: a handover during someone's leave, an error found after a report went out, a colleague missing internal deadlines, a request to skip a required review, credit taken in a meeting, and a colleague contradicting you in front of a client
Responses per scenario4, ranked as a complete ordering rather than rated one at a time. Ranking is used because it forces a trade-off between two options a rating scale would let you call equally good
Explanations shown24 — one for every option, not only for the ones you placed differently. Each carries the kind of response it represents and the competency an employer's key would file it under
Option orderRedrawn from a fresh seed on every run, and independently per scenario, so nothing can be learned from where an option sits on screen
ScoringNone. No total, no percentage, no percentile, no pass mark and no comparison with other visitors — the page reports your ordering beside its own and counts your own placements by type
ClockThere is none, and nothing on the page is measured in milliseconds. Most published situational judgement tests are either untimed or given a limit generous enough that reading speed is not what is being measured
Source for the formatMotowidlo, Dunnette & Carter (1990), An alternative selection procedure: The low-fidelity simulation, Journal of Applied Psychology — the paper that introduced the written scenario as a stand-in for a work sample, and coined the low-fidelity simulation
What leaves the pageNothing. There is no request to any server while you rank, no answer is stored, and the copy button writes your six orderings to your own clipboard as plain text

Frequently asked questions

Is there a right answer to a situational judgement test?

There is a keyed answer, and it belongs to the employer rather than to the world. The key is produced by putting the same scenarios to people already doing that job well, recording how they rank the options, and scoring candidates on how closely their ranking matches that pattern. Change the employer and the pattern moves — a scenario about escalating a risk is keyed one way in an air traffic control tower and another way in a startup where escalating everything is how you become the bottleneck. So the honest answer is that the option that scores best is whatever the panel behind that particular test agreed on, and the only way to prepare for it is to understand the format and what the employer says it values.

Should I answer as myself or as the model employee?

Read the instruction wording first, because it tells you which question is actually being asked. An instrument that asks what you would most likely do is measuring a behavioral tendency and correlates more with personality; one that asks what the best response is measures knowledge of effective action and correlates more with reasoning ability. Answering the second question when the first was asked is the commonest way candidates make themselves look implausible, because the resulting profile is uniformly heroic across every scenario. The safest reading is to take the instruction literally and, where it asks what you would do, answer for the version of yourself that has read the employer's stated values.

Two of the four options look almost identical to me. Does the order matter?

In a real scored test, less than you fear. Most keys weight the extremes far more than the middle, because the panel itself usually agrees strongly about which option is best and which is worst and disagrees about the two in between. A candidate who gets the top and bottom right and flips the middle pair typically loses very little. That is also why this page shows an explanation for all four options rather than only for the ones you placed differently: the argument for second against third is where the actual reasoning lives, and it is the part a mark would hide.

Can these tests be faked?

They can be nudged and they are harder to fake than a personality questionnaire, for a structural reason. To fake a rating scale you only need to know which end is desirable, which is usually obvious. To fake a scenario ranking you need to know which of four plausible professional responses this employer's experienced people preferred, and the whole point of the format is that reasonable people disagree about that. Faking also leaves a signature: an ordering that is heroic in every scenario, always intervening and never waiting, is not what experienced people produce, because half of a working week is knowing which problems are not yours today.

Why will this page not tell me how other visitors answered?

Because it does not know, and inventing the number would be worse than not having it. Showing you a distribution of other people's answers would require storing what visitors chose, and nothing you do on this page leaves the tab — there is no account, no upload, and no request to a server while you rank. A site that shows you a confident percentage after a test with no sign-in has either collected far more than it told you, or has made the figure up. Neither is a good reason to trust the rest of the page.

Are video-based scenarios scored differently from written ones?

The scoring mechanism is the same and what changes is what the format adds and costs. A filmed scenario carries tone, timing and body language that a paragraph cannot, which is the argument for it, and it removes the reading-speed component that a text-heavy set carries for anybody working in a second language. What it adds is production cost, which is why video sets tend to be shorter and to be reused across roles for longer. From a candidate's side the preparation is identical: the question is still what an experienced person in that job would do first.

How much can practice on a generic set like this one actually help?

It can teach you the format and it cannot teach you the key, and being clear about the difference is worth more than any amount of drilling. Knowing that responses are ranked rather than rated, that instruction wording changes the question, that keys weight the extremes, and that the panel behind the key is doing the specific job you applied for — all of that is transferable and none of it is obvious the first time a scenario appears on screen. What is not transferable is any particular ordering, including the six on this page.

Where a situational judgement key comes from, and why a generic one cannot exist

The format began as a cheap substitute for a work sample. Watching somebody actually do a job predicts their performance well and costs an employer a day per candidate, so the proposal was to describe the situation in a paragraph and ask what the candidate would do — a simulation with the fidelity turned down until it fitted on a page. That is the paper the whole family descends from: Motowidlo, Dunnette & Carter (1990), An alternative selection procedure: The low-fidelity simulation, Journal of Applied Psychology. The idea that makes it work is not the scenario, which anybody can write, but the key: the scenarios were put to experienced people in the target job, their rankings were recorded, and candidates were scored on the distance between their ordering and that consensus. Everything that gives the format its predictive power is inside that second step.

There is no general key. A situational judgement test is scored against a consensus of experienced people in that specific job at that specific employer, which is what makes it a valid selection tool and what makes a generic version unscoreable. It follows that the number a generic scenario set could hand you would be the distance between your judgement and one anonymous writer’s, dressed as a measurement. This page therefore does the only thing that is left and does it deliberately: it states an ordering in advance, argues for each placement, names the competency an employer’s key would file that option under, and lets you disagree. The disagreements are the useful part. Where you and the argument part company is usually a place where the right answer genuinely depends on the employer — how much autonomy a junior person is expected to take, whether escalating early reads as prudent or as helpless — and noticing which of those axes you sit on is worth more before an assessment than another set of scenarios would be.

One published finding is worth carrying into the real thing, because it changes what you are being asked. The wording of the response instruction splits the format in two: asked what they would do, candidates produce something that behaves like a personality measure; asked what the best response is, they produce something that behaves more like a knowledge or reasoning measure, and the two versions of the same scenario set do not correlate with the same things — McDaniel, Hartman, Whetzel & Grubb (2007), Situational judgment tests, response instructions, and validity: A meta-analysis, Personnel Psychology sets the difference out across the accumulated studies. Read the instruction before the first scenario, not after the third. If the process you are in also contains a reasoning paper against a clock, the mixed drill on this site finds the question type that costs you the most minutes; if it contains a questionnaire about how you usually work, the work-inventory page covers what those are doing, and the map of assessment families shows how much each of the eight has been found to predict.

This is a measurement exercise, not a clinical assessment. It reports what you did on this page against a stated reference and nothing more — it cannot establish a personality type, a leadership style, or the fit for any particular job. Only a qualified professional, working with more than a browser, can make that judgment.

Where your six orderings are kept

Every number on this page is worked out by JavaScript running in the tab you are reading it in. Your answers, your reaction times and your score are never uploaded, logged or kept — which is also why the test carries on working after you disconnect from the network, and why nothing here can be held back behind an email address.

Two things follow from that here specifically. There is no cross-visitor tally on this page — not because it would be difficult, but because publishing one would mean keeping what you ranked. And a reload loses your six orderings entirely, which is what the copy button exists for.