Free forced-choice demo, 10 balanced pairs
Inside the Caliper test format: ten pairs, two job models
The personality half of this profile does not ask you to rate statements — it puts two statements against each other, both chosen to look equally good to an employer, and makes you take one. This page is free and needs no account: it runs ten such pairs, arranged so that every one of five dispositions meets every other exactly once, and then scores your identical ten answers against two different job models so you can watch the two totals disagree. That disagreement is the answer to whether you passed, and it is why no page outside the employer can give you a number.
- 100% free
- No signup
- 10 pairs
- Complete round-robin
- 2 job models, weights shown
Ten pairs, and both options are good
The personality section of this profile is not a rating scale. It puts statements against each other and makes you pick, and the statements in a pair are matched for how good they look — so the strategy that works on a rating inventory, agreeing with everything flattering, has nothing to grip. What you cannot do is take both.
Below are ten pairs covering five dispositions, arranged so that every disposition meets every other one exactly once. Pick the statement that is more true of you in each. There is no clock, and neither option in any pair is the keyed answer, because there is no key.
How to read a forced-choice block set
Take one of two good options, ten times, then watch two employers disagree about the result.
Pick the statement that is more true of you, not the better one
Both statements in every pair were written to be desirable at work, so there is no better one to find. If a pair feels like a coin toss, the pair is doing its job — the whole reason this format exists is that the strategy which works on a rating scale, agreeing with everything flattering, has nothing to grip on when both options are flattering. Both options are buttons and either can be reached from a keyboard or a phone.
Check the first column of the results table and notice what it cannot do
Your five counts add to ten, because each of the ten pairs gave one point to one disposition. That is not a coincidence to be admired; it is the property that makes a score of this kind incomparable between people. Your highest disposition is only highest relative to your own other four, and the four scales below it were competing for whatever the top one did not take.
Read the two job-model totals against each other rather than separately
The same ten picks are multiplied by two different sets of weights, both printed in full underneath. One model rewards persuading and urgency, the other rewards checking and questioning, and the gap between the two totals is the figure at the top of the panel. Neither number is a verdict on you and neither model is real — they are there so you can see that the number follows from the weights, which are the employer's property and are not published.
Technical specifications
| Blocks | 10 forced-choice pairs, which is exactly the complete round-robin over five dispositions: every pair of dispositions meets once, and each disposition appears in exactly 4 pairs |
|---|---|
| Statements | 20, written for this page, four per disposition, each used in exactly one pair so nothing is asked about twice |
| Matching within a pair | Both statements are written to be plainly desirable at work. A pair where one option is obviously better is not a forced choice, it is a rating scale with extra steps |
| Score total | Fixed at 10 for everybody, before anybody answers. Every pick moves one point, so raising one disposition necessarily lowers the others |
| Forced correlation between scales | −0.25 on average, which is −1 over 4 — a consequence of five numbers summing to a constant and present before any data exists. Not an empirical finding and not a property of these particular items |
| Job models | 2, with all ten weights printed on the page: outbound sales weighted 3 / 3 / 2 / 0 / 0 and clinical-trial data auditing weighted 0 / 0 / 1 / 3 / 3 across the five dispositions |
| Model total range | 0 to 30, being ten picks times a maximum weight of 3. The page reports both totals and the distance between them |
| Percentile or cutoff | Neither, at any point. The real instrument's norms are published only to licensed customers, and the abstract-reasoning half of the profile is not simulated here at all |
Frequently asked questions
What is a good score on the Caliper Profile?
The question has no answer outside the employer, and the demonstration above shows why in a form you can check. A profile is scored against a model of the job rather than against a general population, so the same set of answers produces a strong total against one role's weights and a weak one against another's — the panel gives you both numbers from a single set of picks. That is the design working as intended rather than a flaw: the employer is asking whether you resemble the people who succeed in this particular job, not whether you rank highly on some universal scale. Any site quoting you a target score for this instrument is quoting a number nobody published.
Why does the total always come to the same number?
Because every pair awards exactly one point, so ten pairs award ten points however you answer them. A measure with that property is called ipsative, and the consequence is that the scales stop being independent: with five of them summing to a constant, the average correlation between any two is forced to minus a quarter before a single person has answered anything. This is the mathematical core of the long-standing objection to the format — one person's score on one scale cannot be compared with another person's, because both are statements about internal ranking rather than about quantity.
So is forced choice better or worse than rating statements?
It trades a fakeability problem for a comparability problem, and which trade you prefer depends on what you need the scores for. Removing the obviously desirable answer does make deliberate self-presentation much harder, which is the reason employers buy the format. What it costs is that the raw output cannot be ranked across candidates without further modelling, and the modelling exists but is not what a bar chart in a report is showing. Publishers who use the format properly recover comparable scores by treating each pair as a comparative judgment with an estimated threshold rather than as a point awarded.
Is the abstract-reasoning section on this page?
No, deliberately, and that is worth knowing before you sit the real thing. The published profile pairs its forced-choice personality section with reasoning items — number and letter sequences, figure series — and that half is an entirely different kind of test: it has right answers, it is timed, and practice on it genuinely transfers. Simulating it inside a page about forced choice would confuse the two, so the reasoning drills on this site are kept separate and are linked from the results panel. If your invitation quotes a long completion time, the reasoning half is usually most of it.
Can I tell which disposition a statement belongs to while I am answering?
Usually you can, and it matters less than it feels like it should. Guessing the intended dimension does not tell you which one this employer's model rewards, and that is the piece of information that would actually be worth having — as the two totals in the panel demonstrate, the same pick is an asset under one set of weights and a cost under another. Where guessing does have an effect in a real administration is when a candidate guesses consistently wrong about the target role and shapes a profile away from themselves and away from the job at the same time.
Does answering honestly disadvantage me against people who game it?
The evidence on that runs against the intuition, and it is worth carrying into any assessment. Attempts to correct personality scores for distortion have been tested directly, and they reliably change which candidates would have been hired without improving how well the scores predict later performance — which means the corrections are moving people around for nothing. A forced-choice format adds a second obstacle on top of that: to shape a profile you would need to know which of two equally attractive statements the employer's model rewards, and that information is inside the model rather than inside the item.
How long does the real profile take, and can I pause it?
The invitation itself is the only reliable source on both counts, and it usually states a completion window rather than a per-question limit. What is safe to say is that the personality section is untimed in the sense that individual pairs are not clocked, while the reasoning section is timed, so the total figure in the invitation is not evenly divided between the two halves. The practical advice is to start the reasoning half only when you have a clear run at it, which is the opposite of the advice for the pair-choosing half — that one you can work through at whatever pace you read at.
Where forced choice came from, what it fixed, and what it broke
The idea is a century old and did not start in hiring. Presenting two things and asking which is greater, rather than asking for a rating, is the method psychophysics developed because human beings are far better at comparing than at assigning absolute numbers, and selection testing borrowed it for a specific complaint: on a rating scale an applicant can endorse every attractive statement, and many of them do. Matching the two options in a pair for how attractive they look removes the strategy at its source. There is no desirable end to move toward when both ends are desirable, which is exactly what the ten pairs above are for.
The cost arrived with the fix and is arithmetic rather than empirical. When every block awards a fixed number of points, the totals sum to a constant, and constant-sum scores cannot all rise together — so the scales are forced into negative correlation with each other, at an average of −0.25 for a five-scale instrument, before any person has been measured. Hicks (1970), Some properties of ipsative, normative, and forced-choice normative measures, Psychological Bulletin set this out and it has never been refuted, because there is nothing to refute: a high score means high relative to that candidate’s own remaining scales, and comparing one candidate’s scale against another’s is comparing two internal rankings as though they were quantities. The modern answer is not to abandon the format but to stop reading the raw counts: treating each pair as a comparative judgment with an estimated threshold recovers scores that can be compared across people, which is what Brown & Maydeu-Olivares (2011), Item response modeling of forced-choice questionnaires, Educational and Psychological Measurement made practical. A bar chart in a candidate report is not showing you that model, though, which is why the panel above shows the counts and then tells you what they will not support.
Every one of these publishes its norms only to licensed customers, and several norm by role and by employer rather than against a general population, so a single 'average score' would be wrong even if it were public. The item banks are copyrighted as well: reproducing them would breach the spec's rule that these pages are practice and preparation, never the instrument. The same applies to the mechanical comprehension tests these pages sit next to, which are sold as hiring instruments with their own restricted norms. That is why the two job models on this page are invented and printed in full rather than borrowed and hidden. It also frames the question candidates ask most often, which is whether to shade their answers: correcting personality scores for distortion has been tested against criterion data and it changes who would have been hired without improving the prediction — Christiansen, Goffin, Johnston & Rothstein (1994), Correcting the 16PF for faking: Effects on criterion-related validity and individual hiring decisions, Personnel Psychology put numbers on both halves of that. If the same process has sent you the reasoning half as well, that is the part where an evening is well spent: the mixed drill locates the question type that costs you time, the page on a raw count out of fifty explains where the arithmetic on such a score stops, and the diagram battery covers the version used for technical and utility roles. For the rating-scale format this page is the counterpart to, and a measurement of how far it bends, the inventory page runs the other half of the comparison.
This is a measurement exercise, not a clinical assessment. It reports what you did on this page against a stated reference and nothing more — it cannot establish a personality type, a work style, or your suitability for any particular job. Only a qualified professional, working with more than a browser, can make that judgment.
Not affiliated with, endorsed by, or connected to the owner of Caliper Profile. Caliper Profile is a trademark of its owner and is used here only to name the assessment this page prepares for. The questions on this page are our own: no part of the published test is reproduced, and a score here is not comparable to a score from the real instrument.
Where your ten picks are kept
Every number on this page is worked out by JavaScript running in the tab you are reading it in. Your answers, your reaction times and your score are never uploaded, logged or kept — which is also why the test carries on working after you disconnect from the network, and why nothing here can be held back behind an email address.
One clause specific to this page: both job-model totals are computed in the tab from weights that are written into the page you already downloaded, so there is no request to anything while you answer and no version of your profile exists anywhere to be compared against a real model. A reload clears the ten picks.