Popular test
MBTI
Myers-Briggs Type Indicator
The most famous personality instrument in the world, and the clearest case on this site of popularity and evidence pointing in opposite directions.
Why that grade: The four scales have respectable internal consistency, but the type assignments they produce are unstable on retest and the underlying dimensions are continuous rather than the two categories the model requires.
- Cost
- $59.95 to $99.95
- Time
- 15-25 min
- Items
- 93
- Published
- 1998
Step I (Form M) uses 93 items. Step II (Form Q) uses 144 items and reports 20 facets underneath the four letters.
Publisher
The Myers-Briggs CompanyOrigin
Katharine Cook Briggs and Isabel Briggs Myers, interpreting Carl Jung’s Psychological Types (1921)
What you get free: None from the publisher. The free four-letter tests all over the internet are imitations, not the MBTI.
The verdict
Worth doing once, in a room with other people, for the vocabulary. Not worth using to pick a career, screen a candidate, or explain a relationship. If what drew you to it was the promise of understanding yourself accurately, the Big Five will do that better and for nothing.
Strengths
- A genuinely useful shared vocabulary. Telling a colleague you need processing time before a decision is easier when both of you have a word for it.
- Non-evaluative by design. No type is presented as better, which makes it far less threatening than trait scores in a team setting.
- Step II facets are more informative than the four letters and are the part of the product most worth paying for.
- The four scales do track real variance, so the description you get is rarely wildly wrong, just overstated.
Limitations
- Retake it in a month and there is a meaningful chance you will be told you are someone else, which is disqualifying for anything consequential.
- Forcing a continuous score into one of two boxes discards most of the information, and does the most damage to the people nearest the middle, who are the majority.
- It has no Neuroticism dimension, so it is silent on the trait most connected to wellbeing.
- The Myers-Briggs Company itself states the instrument should not be used for hiring or selection, and it is used that way constantly.
- Almost every free "MBTI test" online is an unlicensed imitation with unknown scoring.
Where the model came from
Carl Jung proposed psychological types in 1921 as a theoretical scheme, not a measurement system. Katharine Cook Briggs and her daughter Isabel Briggs Myers turned it into a questionnaire during and after the Second World War, initially to help place women entering wartime industrial work. Neither had formal training in psychometrics, which is a fact about the instrument’s history rather than an insult: the psychometric refinement came later, from the publisher and from academic critics.
The fourth dichotomy, Judging versus Perceiving, is Myers’ own addition rather than Jung’s. It is also, as it happens, the scale that correlates most cleanly with a Big Five factor.
The dichotomy problem, stated plainly
The type model requires that people cluster into two groups on each scale. They do not. Scores pile up in the middle and thin out toward the edges, the way height does. When you cut a single hump down the middle and name the halves, you create two categories that feel meaningful and are mostly an artefact of where you put the knife.
This is why retest instability is not a minor quality-control issue but a direct consequence of the design. Someone one point above the Thinking-Feeling cutoff is not a Thinker in any stable sense. They are near the middle, and a slightly different mood on a Tuesday will move them across.
What it is still good for
A shared vocabulary has real value even when the categories behind it are soft. Teams that have run an MBTI workshop often communicate better afterwards, and it is not obvious that the improvement requires the model to be true. Being given permission to say "I process this kind of thing internally first" is useful whether or not Introversion is a category.
The honest framing is that the MBTI is a well-designed conversation structure wearing the costume of a measurement instrument. Use the conversation. Do not use the measurement for anything that affects someone’s livelihood.
What it reports back
Results are expressed as discrete types.
- Extraversion / Introversion
- Where attention and energy are directed, outward or inward.
- Sensing / Intuition
- Whether you take in information concretely and sequentially or by pattern and implication.
- Thinking / Feeling
- Whether decisions are weighed by impersonal logic or by impact on people.
- Judging / Perceiving
- Whether you prefer matters settled and planned or kept open and adaptable.
The evidence, with sources
Nothing in this section is stated without an attribution. Where a figure varies by version or sample, the range is given rather than the most flattering number.
Reliability
A substantial share of people who retake the instrument after a few weeks receive a different four-letter type, because scores near a scale midpoint flip category on small changes in responding.
peer reviewedPittenger (2005), Consulting Psychology Journal: Practice and Research, 57(3), 210-221The continuous scores behind the letters are considerably more reliable than the letters themselves, which is an argument for reporting the scores and discarding the types.
peer reviewedMcCrae & Costa (1989), Journal of Personality, 57(1), 17-40
Validity
Four of the MBTI scales correlate substantially with four of the Big Five factors, so the instrument measures real trait variance. It simply measures less of it than the Big Five does, omitting the Neuroticism dimension entirely.
peer reviewedMcCrae & Costa (1989), Journal of Personality, 57(1), 17-40A US National Research Council committee reviewing the instrument found the evidence base did not support the type dichotomies or the career-outcome claims being made for it.
review bodyDruckman & Bjork (eds.), In the Mind’s Eye: Enhancing Human Performance, National Research CouncilScores on the four scales are distributed as a single hump around the middle rather than as two separate clusters, which is the distribution the type model needs and does not get.
peer reviewedPittenger (2005), Consulting Psychology Journal: Practice and Research, 57(3), 210-221
Who it actually fits
- Team workshops where the goal is shared language rather than measurement
- Readers who want a starting vocabulary for talking about differences
- Anyone who already knows their type and wants to know what it is worth
Not a diagnosis