Skip to main content

Editorial standards

How the writing on this site gets made.

The Defaults Research TeamEdited by Ruslan ShaymardanovLast reviewed

Defaults publishes articles about personality science and sells a $2 report scored against published norms. Both rest on the same claim: that what we say is traceable to something you can check. This page describes how the articles are produced, including the parts that are usually left out.

If you find an error on this site, please tell us. What happens next is written out below.

Who writes this

Articles carry the byline The Defaults Research Team. That is the research function of Defaults, and we would rather tell you exactly what it is than let the phrase imply something larger: today it is one person, Ruslan Shaymardanov, working with AI research tooling against the published literature. He is named as editor on every article and is accountable for every word on this site.

There are no other authors, and there are no invented ones. You will not find a credentialed persona here who does not exist. If that ever changes and real contributors join, they will be named individually.

No one on this site is a clinician. The editor holds no psychology license, no therapy credential, and no clinical standing in any jurisdiction, and nothing published here — Defaults included — works as a diagnostic tool, medical advice, or a diagnosis. Where an article touches a topic with a clinical version — depression, ADHD, burnout — it says so, and it points toward a validated instrument or a professional instead of letting a personality score stand in for either.

Where the claims come from

Articles draw on a fixed corpus of peer-reviewed papers identified by DOI — the same corpus that constrains the generated reports, not a separate looser one for marketing copy. You can read what is in it and the full citation list.

We cite primary sources rather than summaries of them. Where a paper is paywalled we link the DOI anyway, so you can see exactly which paper is being leaned on even if you cannot read it for free.

The instrument itself has its own page. How the IPIP-NEO-120 is keyed, whose norms it is scored against, the percentile formula, and the reliability figures are all written out under methodology, with all 120 items published verbatim.

How an article is made

1. We find out what people actually ask. Topics come from real query data — search autocomplete, our own Search Console impressions, and question sites — rather than from guessing what sounds interesting. A topic nobody asks about does not get written.

2. We check what we have already covered. New topics are diffed against the existing corpus so we are adding an answer rather than a near-duplicate of one already on the site.

3. We select the sources first. The papers that can support the article are chosen before the article is written, from the corpus above. This is the step that decides what the piece is allowed to claim.

4. Drafting is AI-assisted. We use AI tooling to synthesise the selected literature into plain language. We are telling you this because it is true and because you would be right to wonder. What the tooling is not permitted to do is supply the evidence: it works from sources chosen in step 3, and a claim it produces without one does not survive step 5.

5. A human checks it against the sources. Before publication the draft is read against the papers it cites, to confirm that every claim is one those papers actually support and that no number has been invented or inflated. This is the step that makes the byline mean something.

What we will not do

No invented statistics. No effect size, correlation or percentage appears without a source that reports it. A plausible-sounding number with no paper behind it is the one failure that would discredit everything else here.

No borrowed authority. The norms we score against come from Johnson (2014) and apply to the IPIP-NEO-120. They are not applied to the adapted HEXACO items, which have no published norms of ours to stand on — a rule enforced in the scoring code itself, not just in editorial intent.

No diagnosis, and no diagnostic language. Trait scores are not symptoms and a percentile is not a condition.

No horoscope grammar. Writing that flatters by staying vague enough to fit anyone is the oldest trick in this field. Statements that would be true of almost everyone are removed rather than kept for their warmth.

No fabricated authors, reviewers or testimonials.

Corrections

If something here is wrong, email hello@knowthydefaults.com with the page and what is wrong with it. A citation that does not support the claim attached to it is the most useful thing you can report.

When a factual error is confirmed we correct the article and move its visible “Updated” date, so the change is legible rather than silent. Where the error was substantive — a wrong number, a misattributed finding — the correction is described in the article rather than quietly patched.

This has already happened. Two of the most-read pages on this site claimed the IPIP-NEO-120 had test–retest reliability around 0.85, footnoted to Johnson (2014). That paper does not report test–retest reliability at all; it reports internal consistency, roughly .86 to .92 across the five domain scales. Both pages were corrected and the citation discipline that let it through was tightened.

What this evidence cannot tell you

Almost all of it is correlational. Personality research overwhelmingly reports associations, not causes. When an article says conscientiousness predicts an outcome, that is a statement about how two measurements move together across many people, not a claim that one produces the other in you.

Effect sizes are population-level. A real, replicated finding can still be nearly useless for predicting one person. Group differences of the size typically found here leave enormous overlap between individuals.

The norms are drawn from one specific group of test-takers. Percentiles compare you to the people in Johnson’s reference sample, which skews toward English-speaking internet users who volunteered to take a personality test — a real reference group, but not a universal human baseline.

Short scales are imprecise. Each of the 30 facets is measured with four items. That is enough to be informative and not enough to be exact — a facet percentile carries meaningful measurement error, which is why the report treats middling facet scores as uninformative rather than narrating them.