Free precision demo, 50 items, 5 checkpoints
Personality assessment for jobs, and why your result moves
Almost every work questionnaire an employer sends is measuring the same five dimensions under whatever names its publisher owns, and the question candidates actually have about them is why two sittings a month apart disagree. That is a question about precision rather than about personality, and precision is something this page can compute from your own answers: rate ten items and it shows the estimate a ten-item questionnaire would have handed you, then keep going to fifty and watch the same estimate move. It is free, there is nothing to sign into, and it produces no type, no label and no percentile — the output is how far the number drifted and how wide the interval around it is.
- 100% free
- No signup
- 50 items
- 5 checkpoints
- No type, no percentile
Answer ten items. Then answer forty more and watch the first answer move.
Almost every work inventory sits on the same five dimensions under whatever names its publisher uses, and this uses them openly. What it will not do is tell you which one you are. After the tenth item it shows you the estimate a ten-item questionnaire would have handed you, and then it keeps going to fifty and shows you what that same estimate does while the items accumulate.
The distance it moves is the point. It is the reason two inventories taken a month apart disagree, and it is a property of questionnaire length rather than of you having changed.
- 50 items, one click each, no clock. Half of them are worded the opposite way round on purpose.
- You can stop at any checkpoint. Stopping early is itself informative: the interval the page prints gets visibly wider.
How to watch a questionnaire estimate settle
Ten items for a first reading, forty more to see what it was worth.
Answer the first ten and read the interval, not the number
After ten items every dimension has two items behind it, which is enough for an estimate and nowhere near enough for a precise one. The panel prints each estimate with a plus-or-minus beside it: that is half a 95% interval computed from your own answers, and at two items per dimension it is wide enough to cover most of the scale. Watch that figure rather than the estimate. It is the honest version of what a ten-item quiz is telling you.
Keep going to fifty and let the estimate wander
The running figures update as items accumulate, and the page records where each one sat at ten, twenty, thirty, forty and fifty. Nothing about you changes over those minutes, so anything the numbers do is the questionnaire rather than the person. Half the items are worded the opposite way round and are scored as six minus your rating, which is why agreeing with everything does not produce a high profile here.
Read the drift column and the split-half figure together
The drift is the distance between the highest and lowest reading each dimension took across the five checkpoints, in scale points on a range four points wide. The split-half figure underneath it comes at the same thing from another direction: your profile from one half of the items against your profile from the other half, stepped up to full length by the formula that exists for that job. A large drift and a low agreement are two views of the same shortage of items.
Technical specifications
| Items | 50, written for this page: 10 per dimension, of which 5 are keyed toward the dimension and 5 away from it and scored as 6 minus the rating |
|---|---|
| Item order | Five blocks of ten. Each block carries one keyed-toward and one keyed-away item for every dimension, and the two directions alternate through the block so no run of five agreeable-sounding statements occurs |
| Checkpoints | 5, at items 10, 20, 30, 40 and 50, so every dimension is re-estimated on 2, 4, 6, 8 and 10 items |
| Interval | A 95% Student-t interval on each dimension's ten keyed ratings, computed on n−1 degrees of freedom rather than a normal approximation, because on ten values the normal interval is meaningfully too narrow |
| Split-half figure | Spearman rank correlation between the five-dimension profile from one half of the items and the profile from the other, stepped to full length by the prophecy formula. Five points, so the figure carries real error and the panel says so beside it |
| Half assignment | Balanced on wording direction: each block sends its keyed-toward item to one half and its keyed-away item to the other, alternating by block. Splitting odd against even by position would have put every keyed-toward item in one half and measured agreement with the questionnaire instead |
| Dimensional model | The five-factor description, as set out in Goldberg (1990), An alternative 'description of personality': The Big-Five factor structure, Journal of Personality and Social Psychology. Dimension names are the scientific ones and match no publisher's proprietary scale labels |
| What is not produced | No type, no letter code, no percentile, no band, no comparison with other visitors and no statement about job fit. The five means exist so the drift and the intervals have something to be about |
Frequently asked questions
Why do I get a different result every time I take one of these?
Mostly because a questionnaire scale estimates a position from a finite number of items, and a finite number of items pins that position down to a range rather than to a point. The demonstration above puts a figure on it without you having to take anything twice: it reports the same estimate at five different questionnaire lengths in one sitting, and the distance it travels is produced by nothing but asking more questions. Two sittings a month apart add mood, wording differences and whatever you were thinking about that morning on top of that. What almost never explains the gap is your personality having changed, which is the explanation people reach for first.
Which five dimensions is my employer's questionnaire actually measuring?
Under the trade names, almost certainly these five or a rearrangement of them, because forty years of factor-analytic work keeps recovering the same structure from independently written item pools. Publishers differ in how they cut them: some report facets underneath each dimension, some combine two of them into a single composite with a proprietary name, and some report a set of scales that do not look like the five at all until you see how they intercorrelate. This is worth knowing for one practical reason — the profile you get from two different vendors will not be identical, but it will not be independent either, so a result that surprises you from one is unlikely to be reversed by the other.
What is a validity scale or a social-desirability scale doing?
Counting how flattering your answers were, using statements that are desirable and that almost nobody can truthfully endorse. The technique dates to a scale published in 1960 built from exactly such items, and a modern inventory usually adds a consistency check on top: pairs of items that say near-identical things in opposite directions, where answering both the same way is arithmetically inconsistent. What happens to the count afterwards is the interesting part, and it is not what candidates assume. Adjusting somebody's scores for how desirably they answered has been tested repeatedly against later job performance and it does not improve the prediction, although it does change who would have been hired.
Should I answer honestly if I think other candidates are not?
The evidence gives an unusually clear answer for this field: shading your answers has not been shown to help you, and the corrections employers apply for it have not been shown to help them. There are two practical reasons to answer as accurately as you can. The consistency checks are looking for the pattern that deliberate shaping produces rather than for the shaping itself, and a profile that is uniformly excellent across every dimension is the specific pattern they notice. And a role you got by describing somebody else is a role that will be measured against that description afterwards, by people who work next to you.
How much does a personality questionnaire actually predict about job performance?
Less than a reasoning test and more than nothing, with conscientiousness the dimension that holds up most consistently across job families. The 1991 meta-analysis that established that is still the reference point, and the coefficients in it are small — small enough that a serious reconsideration published in 2007 argued the field had been over-selling them, particularly where a single dimension is used to screen. What keeps these instruments in hiring is not their standalone strength but that they add something a reasoning test does not, and that they cost almost nothing to administer at volume.
Is a ten-item version just a shorter version of the same measurement?
No, and the panel above shows the difference rather than asserting it. Precision improves with the square root of the item count, so ten items give roughly half the precision of forty rather than a quarter of it — the returns are real but they diminish, which is why serious inventories run to a couple of hundred items and then stop. What a ten-item quiz produces is a reading whose interval covers most of the scale, presented as a point. That is the specific misrepresentation this page exists to make visible, and it is also why a free quiz and a licensed instrument can disagree completely without either being broken.
Can this page tell me my type or which job suits me?
It cannot and it does not try, and the reason is not caution about liability. Nothing here is scored against a population, because the norms for the commercial instruments in this space are published only to their licensees and this site invents none — so there is no distribution to place you in and no percentile to print. Beyond that, a type is a claim about a person that self-report over fifty items cannot support at the precision the claim implies, which is exactly what the drift column demonstrates. The page reports how well fifty items pinned down their own answer, and stops there.
What sits under the trade names, and what the numbers are worth
The five-dimension structure was not designed; it was found, repeatedly, by factoring independently assembled pools of trait words and watching the same five clusters come back — Goldberg (1990), An alternative 'description of personality': The Big-Five factor structure, Journal of Personality and Social Psychology is the paper that made the case in the form the field settled on. That is why a work inventory from one publisher and one from another are not independent measurements even when their scale names share no vocabulary: both are sampling the same space, so their outputs correlate whether or not either vendor says so. It is also why the interesting question about a questionnaire is not which dimensions it covers but how precisely it locates you on them, which is the question the panel above answers with your own answers rather than with a claim.
Precision here comes from a piece of arithmetic that is older than most of the industry. A scale built from more items is more reliable than one built from fewer, and the relationship between the two was published twice in the same year by two authors who arrived at it independently — Spearman (1910), Correlation calculated from faulty data, British Journal of Psychology is one half of that pair, and the formula still carries both names. Its practical consequence is the one candidates never get told: the returns diminish as a square root, so doubling a questionnaire buys considerably less than twice the precision, and a ten-item version of a two-hundred-item instrument is not a short version of the same measurement. It is a different measurement with a much wider interval, reported as a point. Where a real instrument is honest about this it prints a standard error of measurement beside the score, and most candidate-facing reports do not.
No percentile appears anywhere on this page, and the reason is structural rather than squeamish. Every commercial instrument in this space keeps its tables behind a license, and more awkwardly, several of them are normed against the people already doing a particular job at a particular employer rather than against a working population — so even a leaked table would answer a question about one comparison group instead of about you. What is left, once no percentile can be printed, is the evidence about whether any of it predicts anything — and there the picture is modest rather than empty. Conscientiousness is the dimension that holds up most consistently across job families, established in Barrick & Mount (1991), The Big Five personality dimensions and job performance: A meta-analysis, Personnel Psychology, and the coefficients are small enough that Morgeson, Campion, Dipboye, Hollenbeck, Murphy & Schmitt (2007), Reconsidering the use of personality tests in personnel selection contexts, Personnel Psychology argued the field had been overselling them, particularly where one dimension is used as a screen. On the question candidates ask most, the answer is unusually clean: Ones, Viswesvaran & Reiss (1996), Role of social desirability in personality testing for personnel selection: The red herring, Journal of Applied Psychology tested whether correcting scores for socially desirable responding — the count that a scale in the tradition of Crowne & Marlowe (1960), A new scale of social desirability independent of psychopathology, Journal of Consulting Psychology is keeping — improves the prediction of job performance, and it does not, while it does change who would have been hired. The two formats this arrives in have their own pages: a rating inventory can be pushed, and there you can measure by how much, while a forced-choice one removes the obvious answer at the cost of scores that cannot be ranked between people. If the same process also sent you something with right answers on it, that is where an evening of preparation actually returns something — fifty items in twelve minutes and the scenario format are the two most likely to be in the envelope alongside this one, and the behavioral-and-cognitive pair covers the vendor that sends both in one link.
This is a measurement exercise, not a clinical assessment. It reports what you did on this page against a stated reference and nothing more — it cannot establish a personality type, a disorder, or a trait an employer would name. Only a qualified professional, working with more than a browser, can make that judgment.
Not affiliated with, endorsed by, or connected to Hogan Assessment Systems, Inc.. Hogan is a trademark of its owner and is used here only to name the assessment this page prepares for. The questions on this page are our own: no part of the published test is reproduced, and a score here is not comparable to a score from the real instrument. The same holds for every other instrument named on this page, the Caliper Profile among them: each mark belongs to its owner and appears here only to say which instrument is being described.
Where your fifty ratings are kept
Every number on this page is worked out by JavaScript running in the tab you are reading it in. Your answers, your reaction times and your score are never uploaded, logged or kept — which is also why the test carries on working after you disconnect from the network, and why nothing here can be held back behind an email address.
Worth adding for this page in particular: the drift, the intervals and the split-half figure are all computed from your fifty ratings inside the tab, and no part of the calculation needs a population to compare you against — which is convenient, because there is none here. A reload discards all fifty rather than resuming them.