Skip to content
Personality MetricsEvidence first

The guide

How to read a personality test

Enough about how these instruments are built to let you judge any of them yourself, including the ones this site has not reviewed.

The two words that decide everything

Almost every argument about personality tests reduces to two technical terms, and once you have them the arguments become much easier to follow.

Reliability is consistency. If the items meant to measure conscientiousness all correlate with each other, the scale has internal consistency. If you score roughly the same on Tuesday as you did last month, it has retest reliability. Reliability is necessary and it is the easier of the two to achieve. A bathroom scale that reads three kilograms heavy is perfectly reliable.

Validity asks the harder question: does the score mean what the test says it means, and does it predict anything outside itself? This is where most popular instruments come apart. It is entirely possible to build a highly reliable questionnaire that measures nothing anyone cares about, and the field is full of them.

When you see a test described as scientifically validated, the useful follow-up is: validated by whom, published where, and replicated by anyone without a commercial interest? A publisher’s own technical manual is a starting point, not an answer.

Why types keep changing and traits do not

The single most common complaint about personality tests is that the result changed. There is a specific and well-understood reason for this, and it is worth understanding because it tells you which instruments to trust.

Personality traits are distributed as a single hump. Most people are somewhere in the middle on most things, and the extremes thin out, the way height does. There are not two kinds of people, tall and short, with a gap between them.

A type-based instrument takes that single hump and cuts it down the middle, then gives the two halves names. If you land one point above the cutoff you are told you are a Thinker; one point below and you are a Feeler. Since most people are near the middle, most people are near a cutoff, and a slightly different mood on a different day moves them across it.

Trait instruments avoid this by reporting where you sit on the scale rather than which side of a line. The result is less satisfying and considerably more stable. This is the whole reason the Big Five is the research standard and the MBTI is not, despite the MBTI being far more famous.

Reading your own results without overreading them

Three habits will get you most of the value available from any instrument, and protect you from most of the damage.

Look at the facets, not just the headline. Conscientiousness contains both orderliness and self-discipline, and plenty of people are high on one and low on the other. The domain score averages that tension away, and the tension is the interesting part.

Notice how close to the middle you are. A result of 51 percent on some axis is telling you the instrument could barely separate you. Treat that label as noise. Any decent report shows you this number, and almost nobody reads it.

Treat the result as a description, not a constraint. Trait scores describe where you have been sitting under ordinary conditions. Mean levels shift predictably with age, and deliberate effort moves them too. Nothing in the evidence supports reading a score as a ceiling.

The Barnum problem

Personality descriptions are unusually good at feeling accurate. This is partly because they are often true, and partly because a well-written profile is constructed from statements that almost anyone would accept about themselves. The demonstration is old and reliable: give a room full of people the same generic profile, describe it as personalised, and most of them will rate it as highly accurate.

This is not an argument that personality testing is worthless. It is an argument that your sense of recognition, on its own, is not evidence that an instrument is measuring anything. That is what the reliability and validity literature is for, and it is why this site insists on citations for every figure it reports.

When a personality test is the wrong tool

There are three situations where the honest answer is that you want something else entirely.

Hiring and selection. Self-report inventories are easy to distort when a job is at stake, and several publishers explicitly state their instruments must not be used this way. Structured interviews and work-sample tests have better evidence.

Clinical concerns. No instrument on this site can diagnose anything. If you are worried about anxiety, depression, burnout, or anything else affecting your health, that is a conversation with a doctor or a therapist, not a quiz.

Choosing a career from scratch. Personality predicts job satisfaction more than job performance. If the question is what you would be good at, aptitude measurement has a stronger evidence base than personality measurement does.

A reasonable path through all of this

If you want one recommendation: take the BFI-2. It is free, it takes about ten minutes, it has published psychometrics, and it will give you scores rather than a label. If that interests you, follow it with HEXACO for the dimension the Big Five leaves out.

Then, if you like, take the famous ones for the vocabulary and the conversation, knowing exactly what you are getting. There is nothing wrong with enjoying an instrument that does not hold up, as long as nobody is making a decision about your life with it.

Common questions

Frequently asked

What does reliability mean for a personality test?
Reliability is consistency. Internal consistency asks whether items meant to measure the same thing actually correlate with each other. Test-retest reliability asks whether you get the same answer when you take it again. A test can be highly reliable and still measure the wrong thing, which is why validity is the separate and harder question.
What does validity mean for a personality test?
Validity asks whether the instrument measures what it claims to and whether the score predicts anything outside the test itself. It is harder to establish than reliability and is where most popular instruments fall down. A test with good reliability and poor validity is a precise measurement of nothing in particular.
Why did I get a different result the second time?
Almost always because the instrument sorts a continuous score into categories. Most people sit near the middle of most traits, so a small change in mood or wording moves them across a boundary and changes the label entirely. It is a property of the type model, not a fault in you.
Can a personality test diagnose a mental health condition?
No. Every instrument reviewed on this site is a self-report or performance measure for reflection. Clinical assessment uses different instruments, administered and interpreted by qualified professionals, alongside history and interview. If you are worried about your mental health, speak to a doctor rather than taking a quiz.
Should employers use personality tests for hiring?
Self-report personality inventories are easy to distort when something is riding on the answer, and several publishers explicitly state their instruments should not be used for selection. Structured interviews and work-sample tests have stronger evidence for predicting job performance.

If you are struggling right now

Nothing on this site is a clinical assessment or a substitute for professional care. If pressure is affecting your sleep, your health, or your relationships, talk to a doctor or a qualified mental health professional.