Someone posts their results and asks what they mean. The numbers look like percentages, so they get read like percentages: 98 in Openness sounds like being almost entirely open, and 11 in Conscientiousness sounds like having almost none. Neither reading is right, and the gap between them and the truth is where most misinterpretation of these tests begins.
A percentile is a position, not an amount
If your Openness score sits at the 80th percentile, roughly 80% of people in the comparison sample scored lower than you did. It says nothing about how much Openness you have in any absolute sense, because the Big Five has no absolute scale. There is no zero point and no natural unit. The only thing the instrument can do is place you relative to other people who answered the same items.
That has an immediate consequence. A score of 11 does not describe someone with almost no Conscientiousness — such a person could not hold a job or run a household. It describes someone less conscientious than about 89% of the reference sample, which is a far more modest claim, and one that says nothing about whether they are conscientious enough for any particular purpose.
Which people you are being compared to
A percentile is only as meaningful as the group behind it. Ours come from Johnson's published norm tables for the IPIP-NEO-1201, split by sex and age band, so your number is a position among people of roughly your own demographic rather than among everyone alive.
That sample is a real group of people and it is not humanity. It skews toward English-speaking internet users who chose to take a personality test, which is not a random draw. The effect is largest at the extremes: a 95th-percentile score means 95% of that sample scored lower, and a different sample would move the number. We publish the norm tables we score against so the comparison group is inspectable rather than implied.
How much precision the number actually has
This is the part most tests leave out, and it changes how you should use the result.
The IPIP-NEO-120's domain scales have internal consistency of roughly .86 to .921. That is good, and it is not the same as exact. At that reliability, a score sitting at the 50th percentile carries a 95% interval running from about the 27th to the 73rd. That single fact is the most useful thing on this page: two people whose reported scores differ by twenty points near the middle may not differ at all.
Facet scores are less precise again. Each of the 30 facets is measured with four items — enough to be suggestive, not enough to be exact. One changed answer can move a facet by six to seventeen percentile points.
The rule that follows: treat scores near the middle as uninformative and pay attention to the ones near the ends. A score at the 55th percentile is not a finding. A score at the 90th is.
What the number does not tell you
It is not a measure of how much of a trait is good to have. High Conscientiousness predicts job performance across nearly every occupation studied2 and carries real costs in rigidity and slow adaptation. There is no percentile at which a trait becomes correct.
It is also not a category. The Big Five measures continuous dimensions, and data-driven attempts to carve them into discrete personality types recover only a small number of loose clusters rather than clean types3. Scoring at the 70th percentile on Extraversion does not put you in a group; the person at the 65th is barely distinguishable from you.
Reading a whole profile
The question people actually ask is what the pattern means, and that is the right question, because traits interact. High Conscientiousness with high Neuroticism produces a different life from high Conscientiousness with low Neuroticism, in ways neither score predicts alone.
The honest procedure:
- Ignore anything between roughly the 40th and 60th percentiles. That is measurement error wearing a number.
- Find your two most extreme scores. Those are the ones that reliably say something.
- Read what those two do together, rather than reading each separately and adding them up.
The last step is where most self-interpretation goes wrong, and it is the one a bare list of five numbers cannot help you with.
Screen one trait free (3 min) → or take the full Big Five test (12 min, $2, report included) →
Footnotes
-
Johnson, J. A. (2014). Measuring thirty facets of the Five Factor Model with a 120-item public domain inventory: Development of the IPIP-NEO-120. Journal of Research in Personality, 51, 78–89. https://doi.org/10.1016/j.jrp.2014.05.003 ↩ ↩2
-
Barrick, M. R., & Mount, M. K. (1991). The Big Five personality dimensions and job performance: A meta-analysis. Personnel Psychology, 44(1), 1–26. https://doi.org/10.1111/j.1744-6570.1991.tb00688.x ↩
-
Gerlach, M., Farb, B., Revelle, W., & Amaral, L. A. N. (2018). A robust data-driven approach identifies four personality types across four large data sets. Nature Human Behaviour, 2, 735–742. https://doi.org/10.1038/s41562-018-0419-z ↩