Research instrument
The Big Five
Five-Factor Model (OCEAN)
The consensus trait model in academic personality psychology, and the benchmark every other instrument on this site is measured against.
Why that grade: Replicated across languages, measurement methods, and observer reports, with trait-outcome associations that survive large-scale replication attempts.
- Cost
- Free
- Time
- 20-30 min
- Items
- 120
- Published
- 2014
The IPIP-NEO-120 is the common free form. A 300-item long form and a 60-item short form also exist, and the commercial NEO-PI-3 uses 240 items.
Origin
Model developed across decades of lexical research; the free IPIP-NEO-120 form was developed by John A. Johnson
What you get free: The IPIP item pool is in the public domain. Full domain and facet scores, with no account and no upsell.
The verdict
If you take exactly one personality assessment in your life, take this one. It is free, it is the model the research literature is written in, and it is honest enough to give you a position on a scale rather than a flattering label. It is also the least entertaining option on this site, which is the trade you are making.
Strengths
- The only widely available model whose structure has been recovered independently across languages, raters, and decades.
- Reports continuous scores, so it can express "somewhat above average" instead of forcing you into a bucket.
- Free, public-domain item pools mean no paywall stands between you and a full facet-level profile.
- Facet-level scores show within-trait disagreement that a single domain score hides, such as being orderly but not self-disciplined.
Limitations
- The output is genuinely less fun than a four-letter type, and percentile scores relative to a norm group are easy to misread as grades.
- Self-report is distortable. Anyone with a reason to look good can look good, which is why this is a poor hiring instrument despite its research pedigree.
- Openness in particular does not replicate cleanly in every language and culture studied, so cross-cultural comparisons deserve caution.
- It describes traits without explaining them. It will tell you that you score low on Conscientiousness and nothing at all about why.
What it actually measures
The Big Five is not a test. It is a finding: that when you take the personality-descriptive words a language contains and ask enough people to rate themselves on all of them, the ratings collapse into roughly five clusters. That result was arrived at from the vocabulary up rather than designed from a theory down, which is precisely why it carries more weight than models that began with an idea about human nature and went looking for support.
What you take is an inventory built to measure those five clusters. The IPIP-NEO-120 is the most useful free one: 120 statements, scored into five domains and thirty facets. The commercial NEO-PI-3 measures the same territory with more polish and a price attached.
What the evidence supports, and what it does not
The evidence supports the structure and it supports the predictions. Five broad factors keep reappearing across languages and rating methods, and scores on them forecast outcomes that matter, at effect sizes that survive replication. That combination is rare enough in this field to be the whole reason the model is the benchmark.
The evidence does not support treating a percentile as a verdict. A person at the 70th percentile on Extraversion is not a different kind of creature from one at the 55th; the two overlap in nearly every behaviour you could name. Nor does the evidence support using these scores to choose between job candidates, because self-report inventories are trivially easy to game when something is riding on the answer.
Reading your own results without overreading them
Look at the facets before the domains. Conscientiousness contains both orderliness and self-discipline, and plenty of people are high on one and low on the other. The domain score averages that tension away; the facet scores are where the useful, uncomfortable detail lives.
Then treat the whole profile as a description of tendencies under ordinary conditions, not a constraint. Rank-order stability is high, but mean levels shift with age in predictable directions, and deliberate effort moves them too. The score describes where you have been sitting, not where you are required to stay.
What it reports back
Results are expressed as continuous traits.
- Openness to Experience
- Appetite for novelty, abstraction, art, and unfamiliar ideas. The most contested of the five, and the one whose meaning shifts most across cultures.
- Conscientiousness
- Organisation, persistence, impulse control, and follow-through. The single best trait predictor of job performance across occupations.
- Extraversion
- Sociability, assertiveness, and sensitivity to reward. Not the same thing as social skill, and not the opposite of shyness.
- Agreeableness
- Trust, cooperation, and concern for others. High scores predict better relationships and, in some settings, lower earnings.
- Neuroticism
- Frequency and intensity of negative emotion. Often reported inverted as Emotional Stability, which measures the same thing with a friendlier label.
The evidence, with sources
Nothing in this section is stated without an attribution. Where a figure varies by version or sample, the range is given rather than the most flattering number.
Reliability
Johnson reports internal consistencies for the five IPIP-NEO-120 domain scales in the .80s, with the shorter four-item facet scales necessarily lower.
peer reviewedJohnson (2014), Journal of Research in Personality, 51, 78-89Big Five domain scores are among the more stable self-report measures in psychology, with rank-order stability rising through adulthood and peaking in middle age rather than being fixed from birth.
meta analysisRoberts & DelVecchio (2000), Psychological Bulletin, 126(1), 3-25
Validity
Big Five traits predict mortality, divorce, and occupational attainment at magnitudes comparable to socioeconomic status and cognitive ability.
meta analysisRoberts, Kuncel, Shiner, Caspi & Goldberg (2007), Perspectives on Psychological Science, 2(4), 313-345In a preregistered replication of 78 previously reported trait-outcome associations, most replicated in direction and rough magnitude, which is a markedly better record than much of the surrounding literature.
peer reviewedSoto (2019), Psychological Science, 30(5), 711-727Conscientiousness was the one trait to predict performance across all five occupational groups studied, with estimated true-score correlations around .22 averaged across criteria. That is a real and useful effect, and it is nothing like the "87 percent accuracy" figures that circulate in popular writing.
meta analysisBarrick & Mount (1991), Personnel Psychology, 44(1), 1-26; 117 studies, 162 samples, N = 23,994A 2022 reanalysis correcting for systematic overcorrection of range restriction revised the operational validity of conscientiousness for job performance downward, from roughly .26 to roughly .19. The trait still predicts; the older figures were too generous.
peer reviewedSackett, Zhang, Berry & Lievens (2022), Journal of Applied Psychology, 107(11), 2040-2068The free public-domain IPIP scales converge closely with the commercial NEO-PI-R they were built to mirror, at a mean correlation of about .73 across the thirty facet scales, rising to about .94 once corrected for unreliability. Paying does not buy you a materially different measurement.
peer reviewedGoldberg (1999), Personality Psychology in Europe, 7, 7-28; IPIP validity documentation
Who it actually fits
- Anyone who wants the most defensible picture of their personality available for free
- Readers who intend to compare their results against published research
- Students and researchers who need an instrument they can cite
Not a diagnosis