Skip to main content

Why Two Personality Tests Give You Different Results

The Defaults Research TeamEdited by Ruslan ShaymardanovPublished 6 min read

Two Big Five tests can report different percentiles for the same person because they use different items, different scoring, and above all different comparison samples. A percentile is a position within one specific reference group, so changing the group changes the number without anything about you changing.

You take one Big Five test and score 70 on Conscientiousness. You take another a week later and score 48. Nothing about you changed in a week, so one of them must be wrong.

Usually neither is. Four separate things differ between any two personality tests, and each of them moves the number on its own.

The comparison group is different

This is the largest source of disagreement and the least visible, because most tests do not tell you what it is.

A percentile is not a property of your answers. It is a position within a reference sample. Score the identical raw answers against two different samples and you get two different percentiles, with nothing about you having changed. If one test norms against a general-population sample and another against university students, the second will report you as less conscientious, because students are on average lower on Conscientiousness than the adult population.

Many free tests do not norm at all. They convert your raw score to a percentage of the maximum possible — answer every item at the top of the scale and you get 100% — and then display it in a way indistinguishable from a percentile. Those two numbers mean entirely different things, and only one of them is a comparison with other people. We publish the norm tables we use for exactly this reason.

The items are different

The Big Five is a model, not an instrument. Many instruments measure it, and they do not contain the same questions.

The IPIP-NEO-120 uses 120 items covering 30 facets1. The BFI-2 uses 60 items across 15 facets2. Goldberg's original transparent-format markers are shorter again3. Longer instruments measure more precisely and sample the trait more broadly, so a short test may be capturing a narrower slice of the same construct.

The facet coverage matters more than the length. A Conscientiousness scale weighted toward Orderliness will score a tidy, unpunctual person higher than one weighted toward Dutifulness. Both are legitimately measuring Conscientiousness. They are measuring different parts of it.

Measurement error is larger than people expect

Even the same test, taken twice, will not return the identical number. At a reliability of .90, the 95% interval around a mid-range score spans roughly the 27th to the 73rd percentile.

That single fact accounts for a large share of the disagreements people notice. A gap of twenty percentile points between two tests, near the middle of the range, is entirely consistent with both being correct and your true score being one value.

Gaps at the extremes are more meaningful. Scoring 95 on one instrument and 40 on another is not explained by measurement error, and points to a genuine difference in what the two scales contain.

Your answers are different

Self-report responds to how you have been recently. The question is what your typical behaviour is; the available evidence when you answer is mostly the last few weeks.

Take a test during a difficult stretch at work and Neuroticism items are easier to endorse — not because you are lying, but because the instances are more available. This is why retaking after a major life change produces larger shifts than retaking after a quiet month.

Which one to believe

A reasonable ranking, when two results disagree:

  • Prefer the test that tells you which sample it compared you against. If it does not say, it may not have compared you to anyone.
  • Prefer the longer instrument. More items means more precision, and facet-level results tell you where a domain score came from.
  • Prefer the one you answered less carefully. Deliberating over items pulls answers toward how you would like to be; quick responses to typical-behaviour questions tend to be more accurate.
  • Compare the facets, not the domains. If two tests disagree at the domain level, the facet breakdown usually shows exactly which part of the trait each one weighted.

And if both results are near the middle, the disagreement is not telling you anything. Two uninformative numbers can differ without either being wrong.

Screen one trait free (3 min) → or take the full Big Five test (12 min, $2, report included) →


Footnotes

  1. Johnson, J. A. (2014). Measuring thirty facets of the Five Factor Model with a 120-item public domain inventory: Development of the IPIP-NEO-120. Journal of Research in Personality, 51, 78–89. https://doi.org/10.1016/j.jrp.2014.05.003

  2. Soto, C. J., & John, O. P. (2017). The next Big Five Inventory (BFI-2). Journal of Personality and Social Psychology, 113(1), 117–143. https://doi.org/10.1037/pspp0000096

  3. Goldberg, L. R. (1992). The development of markers for the Big-Five factor structure. Psychological Assessment, 4(1), 26–42. https://doi.org/10.1037/1040-3590.4.1.26