Skip to main content

Built on measured traits, not type lore.

The Defaults Research TeamEdited by Ruslan ShaymardanovLast reviewed

Defaults uses the 120-item IPIP-NEO-120 Big Five inventory. The Big Five is a trait model: it estimates continuous scores across Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism rather than assigning a fixed type.

Big Five results are true norm-referenced percentiles — your raw answers scored against published adult norm tables, not rescaled against themselves. The reference sample and the exact formula are stated in full below, with the numbers this site actually computes against.

Reports use cited research as constraints. The system is instructed to cite only from a curated corpus of Big Five, IPIP, outcome, cluster, and intervention papers, with DOI links shown in the report.

The instrument

120 items, answered on a 1–5 agreement scale. These are the verbatim IPIP-NEO-120 items with Johnson’s original keying — not paraphrases, not a shortened adaptation. The inventory is public domain, which is why we can use the items and the matching norm tables together without licensing drift.

You do not have to take “verbatim” on faith. We publish the full IPIP-NEO-120 item list — all 120 questions in Johnson’s original order, grouped by domain and facet, with each item’s keying direction and the raw-sum scoring key beside it. Read them before you answer them if you like — reading an item tells you what it measures, not where you land, and a test that has to keep its questions hidden is telling on itself.

The 120 items resolve to five domains of 24 items each, and each domain to six facets of 4 items each. Scores are computed on the raw sum scale Johnson’s own scoring program uses: 24–120 for a domain, 4–20 for a facet. Reverse-keyed items are flipped before summing.

That item list is the scoring layout, and the test is not administered in it. Items are asked in a rotation across the five domains — a Neuroticism item, then Extraversion, Openness, Agreeableness, Conscientiousness, then the next facet of each — which is the arrangement Johnson’s own inventory presents them in. Two reasons, and neither is cosmetic. Rotating means wherever you stop, you have answered a balanced spread rather than one trait in depth, and the four items of any facet sit twenty-five questions apart instead of back to back, which is what keeps an answer to one from priming the next. It also means nobody is handed a block of consecutive questions about low mood: in the grouped order all 24 Neuroticism items come first, four of them in a row about feeling down, and we are not willing to open a stranger’s test that way. The Depression and Vulnerability facets each contribute exactly one item to the opening thirty, and it is the reverse-keyed one. If you tell us at the start what prompted you to take this, the domain closest to that leads the rotation, so the questions nearest your reason come first.

None of this touches measurement. Every one of the 120 items is asked, once, with its own keying, and answers are recorded against item ids — so the order changes what you are asked when, and nothing about what is scored or what it is scored against.

Norm source and reference sample

The means and standard deviations are the normative groups embedded in Dr. John A. Johnson’s public-domain IPIP-NEO scoring program — eight groups, crossing sex (male / female) with four age bands (under 21, 21–40, 41–60, over 60).

Johnson, J. A. (2014). Measuring thirty facets of the Five Factor Model with a 120-item public domain inventory: Development of the IPIP-NEO-120. Journal of Research in Personality, 51, 78–89. doi.org/10.1016/j.jrp.2014.05.003

We do not ask for your sex or age, so the default reference group is the element-wise average of the male and female 21–40 tables — the same neutral combination Johnson’s program applies for respondents who do not identify as male or female, and the age band that matches this product’s audience. These are the exact domain figures your percentiles are computed against:

Raw-sum norms (24–120 scale) — combined men & women between 21 and 40 years old
DomainMeanSD
Openness87.3812.40
Conscientiousness86.5314.07
Extraversion79.8414.93
Agreeableness88.0612.25
Neuroticism69.5616.32

Facet-level norms use the same eight groups on the 4–20 scale. Within every group, the six facet means sum to their domain mean (within rounding) — we check this as a transcription guard.

The percentile formula

Your raw sum is converted to a z-score against the reference group, then mapped through the standard normal CDF:

percentile = Φ( (raw − mean) / SD ) × 100

Johnson’s program approximates Φ with a cubic; we evaluate it directly via the error function, which is more precise than the integer we end up reporting. Results are rounded to a whole number and clamped to 1–99. That clamp is deliberate: no real reference distribution justifies calling a person the 0th or 100th percentile, and a test that shows you a 0 is telling you about its own arithmetic, not about you.

This matters more than it sounds. Most free Big Five tests rescale your answer average onto a 0–100 bar and label it a percentile. It isn’t one — it is just your own answers stretched to fit, with no reference group involved. A “28th percentile” here means you scored higher than roughly 28% of the adult comparison group. That is a claim that can be checked.

Reliability of the instrument

Reliability is not one number. The figure Johnson’s validation study establishes for the IPIP-NEO-120 is internal consistency, reported as coefficient alpha: how consistently the items within a scale measure the same underlying thing. Across the five domain scales it runs from about 0.86 to 0.92 — Agreeableness at the lower end, Neuroticism at the upper end, with Openness, Extraversion and Conscientiousness in between.

A second statistic, test–retest reliability, asks a different question: whether the same person gets the same score weeks or months later. Johnson (2014) does not report it for this instrument, so we do not quote a number for it anywhere on this site. If you see a test–retest figure for the IPIP-NEO-120 attributed to that paper, it is a misattribution — the 0.85–0.92 range that circulates is the alpha figure above, wearing the wrong label.

Facet scales are built from four items each, so their alphas are necessarily lower than the domain scales they roll up into. That is arithmetic, not a defect — a shorter scale has less room to agree with itself — but it is the reason a facet score deserves a lighter touch than a domain score, and why the reports lead with domains.

Source: Johnson, J. A. (2014). Measuring thirty facets of the Five Factor Model with a 120-item public domain inventory. Journal of Research in Personality, 51, 78–89.

Why HEXACO is deliberately not normed

Published HEXACO-60 norms exist (Ashton & Lee, 2009). We do not apply them, because our shorter HEXACO assessment uses adapted items rather than the published HEXACO-60 item set, and applying one instrument’s norms to another’s items would reintroduce exactly the problem the Big Five scoring above was built to fix.

So HEXACO results are shown as a score out of 100 — how strongly your own answers leaned on each trait — and are never labelled a percentile. If we later license the verbatim items or run a calibration study, that changes. Until then, the honest label is the one on the page.

The same test is free elsewhere. What am I paying for?

It is free elsewhere, and we would rather tell you now than have you find out after paying. The 120 questions are public domain. The items are published at ipip.ori.org, the scoring tables come from Johnson's 2014 paper, and other sites run the same questionnaire at no charge.

What those sites hand you is a page of numbers. If a page of numbers is what you came for, go and take the free one. It is the same questionnaire and we have no quarrel with it.

We do not sell the numbers. We sell what they mean for your week — where the pattern pays off, where it costs you, and which parts of your life it has been quietly running. Written in plain words, put in the order of what it actually touches, and yours to keep. The measured results come with it.

That is also why you do not see a score sheet here before you pay. The score sheet is the part you can already get for free.

How AI is used in your report

The free baseline (your five-domain scores and one anchor insight per domain) is entirely template-driven. No large-language-model is involved in generating it.

The paid report and the chat use Google’s Gemini models. Specific model versions are pinned in our code and updated only through a reviewed change — never silently. The generation step turns your facet scores into the six life-domain narratives; the chat answers your questions with your full facet profile included as context, so the advice is tuned to how you’re wired rather than generic.

What the AI can do: translate your scores into plain-language patterns, suggest scripts, surface trade-offs your profile suggests, answer questions about how to handle situations given your defaults.

What the AI can’t do: diagnose anything, replace a therapist, replace a doctor, replace a lawyer, or know things about you it wasn’t given. It doesn’t have access to the internet, your other accounts, or other users’ data. It can be wrong. Treat its output as a thoughtful starting point, not a verdict.

Sensitive topics

If a conversation touches suicide, self-harm, abuse, or a serious medical or legal situation, the assistant will append this message and step back from offering further advice:

This isn’t therapy and I’m not the right tool for what you’re describing. Please contact the appropriate emergency services or a qualified professional in your location.

We also run an automated review over conversations that get flagged this way. If a human review determines a response may not have been appropriate, we will notify you and surface the same safety message — without quoting the conversation in the email. This review is operated by a small team under the same access controls as everything else on your account.

Your data

Your assessment responses, generated reports, and chat history are kept for as long as your account exists, so you can come back to them. You can clear your chat history at any time from your account settings, and you can export or delete your account on the same page (deletion is a hard delete, cascading through all of your data within 30 days of the request).

The report and chat are generated by Google’s Gemini models, which process your inputs on Google infrastructure in the United States. We do not let Google use your conversations to train or improve its models — calls run on the paid Gemini API, where inputs and outputs are not used for model training.