The Myers-Briggs Type Indicator measures self-reported preferences across four dichotomies, sorting people into one of 16 personality types based on Jungian theory, yet research shows it lacks strong test-retest reliability and predictive validity, making licensed therapy a more reliable path to genuine self-understanding.
Ever taken the Myers Briggs twice and landed on a different type each time? That's not a fluke, it's built into how the test works. Here's what those four letters really measure, what they don't, and why your need for self-understanding deserves more than a label.
What the Myers-Briggs Type Indicator actually measures
The Myers-Briggs Type Indicator (MBTI) is a self-report questionnaire. You answer a series of forced-choice items, meaning you pick between two statements that each describe a way of thinking or behaving, and there is no option to say both apply or neither does. Nothing about the process involves an outside observer rating you or a clinician making a judgment. It is closer to a structured survey of your own preferences than to any kind of assessment a mental health professional would use to evaluate a condition.
What does the Myers-Briggs Type Indicator measure?
The MBTI sorts your answers into four separate dichotomies, each framed as a pair of opposite preferences. These are extraversion versus introversion, sensing versus intuition, thinking versus feeling, and judging versus perceiving. Each dichotomy is meant to capture a general tendency, not a fixed trait you either have or lack entirely. Put together, the four letters you land on make up your reported type.
The four dichotomies and how they are scored
Your raw score on each dichotomy actually falls somewhere on a continuous scale, meaning most people land somewhere in the middle rather than at either extreme. The instrument then converts that continuous score into a single letter, so someone who scores nearly 50-50 between thinking and feeling gets sorted into the same category as someone who scores heavily toward one side. That conversion is where a lot of the disagreement about the tool starts, though the mechanics of why retesting produces different results are a separate question. The letter you receive tells you which side of the midpoint you fell on, not how strongly.
What a four-letter type is meant to describe
Combining the four letters produces one of 16 personality types, each given a label like ENFP or ISTJ. These types are described as preferences for how you direct energy, take in information, make decisions, and organize your life, not as abilities, skills, or guarantees of how well you will perform at anything. The MBTI does not claim to measure mental health, pathology, competence, or job success. That distinction matters: a type describes a preference, while something like the diagnostic criteria behind personality disorders describes a pattern of functioning that causes real distress or impairment, and the two are not measuring the same kind of thing at all.
Type dynamics and the dominant function claim
Beyond the four letters, MBTI theory includes the idea of type dynamics, which proposes that each type relies on an internal hierarchy of four mental functions, with one function operating as dominant and the others playing supporting roles. This dominant function claim suggests that your type is not just four independent preferences but a structured system with a lead process guiding the rest. It is a more elaborate layer than most people encounter when they first take the test, and it moves well past simple description into a theory about how the mind is organized internally.
Where the framework came from, and why that matters
The MBTI history begins not with a questionnaire but with a book. Carl Jung published his theory of psychological types in 1921, describing patterns he had observed in patients and in himself: whether a person’s energy orients outward or inward, and how they take in and judge information. Carl Jung psychological types was a work of philosophy and clinical interpretation, built from case observation rather than measurement. Jung was not trying to build a test. He was trying to describe a structure he believed he saw in the mind.
Katharine Cook Briggs read Jung’s work and became fascinated by it. Her daughter, Isabel Briggs Myers, later turned that fascination into a questionnaire meant to sort people into Jung’s categories. Neither woman had formal training in psychometrics, the discipline concerned with how psychological instruments are constructed and validated. They worked largely outside academic psychology, refining the tool through decades of independent effort rather than university research.
The instrument found its footing in offices rather than clinics. Businesses adopted it for hiring, team building, and career guidance long before psychology departments took much interest in it. That path mattered, because it meant the tool spread and calcified as a workplace habit while academic personality science moved somewhere else. Researchers in that period increasingly favored factor-analytic models, built by statistically grouping traits that actually correlate across large samples, a different starting point than Jung’s typology.
None of this makes the instrument automatically wrong. Origin is not the same as evidence. It does explain a real gap: the theory existed first, built by observation, and the questions about whether it holds up came only after it had already spread far beyond where it started.
Why people get a different type on retest
A lot of people take the Myers-Briggs more than once and land on a different set of letters the second time. That is not random bad luck. It comes from how the test is built and how it turns a number into a label.
What reliability means for a personality test
Reliability means an instrument gives a stable answer when nothing about the person has changed. If you take a test on Monday and again on Friday, and you have not changed in any meaningful way, a reliable test should hand back roughly the same result. Mbti test-retest reliability is the specific question of whether the same person gets the same type letters back on a second try. When the answer is often no, the problem lies in the test itself, not in the person’s stability.
The midpoint problem, explained
Each dichotomy on the Myers-Briggs, thinking versus feeling for example, is scored on a continuous scale before it gets chopped into a single letter. A dimensional score would report something like slightly toward thinking. A categorical label reports T with the same confidence it reports a strong, lopsided preference. People who land near the middle of that scale are the ones most likely to flip letters, and midpoint scores are common rather than rare. A flip changes the label, the type description, and any advice attached to it, even though the underlying score barely moved.
What the retest studies looked at
Reliability research on the instrument has looked directly at how often people change type across repeated testing. Work described in cautionary comments regarding the Myers-Briggs Type Indicator by Pittenger examined this pattern, as did research by McCarley and Carskadon on how consistently respondents landed in the same category across retest intervals. A related psychometric analysis of the MBTI’s scoring limitations ties this instability back to the same source: continuous scores forced into categories at a fixed cut point. Mood, recent events, and the setting in which someone takes the test can all nudge a borderline answer just enough to cross that line, which is often the honest answer to why does my mbti type change between sittings.
Validity, and what the test can and cannot predict
Reliability asks whether you get the same answer twice. Validity asks something different: does the instrument measure what it claims to measure, and does that measurement tell you anything useful. This is the more serious question for the Myers-Briggs, and it is the one where mbti validity runs into trouble. A description can feel exactly right and still fail this test, because accuracy of feeling and accuracy of measurement are not the same thing.
Do the categories describe real divisions between people?
The theory claims that people fall into one of two camps on each dichotomy: extraversion or introversion, thinking or feeling, and so on. If that were true, you would expect trait scores to cluster into two separate humps, one for each side. Instead, research on personality trait structure finds that traits distribute the way height or blood pressure does, bunched around the middle with fewer people at the extremes. That pattern undercuts the idea of a genuine type category and points instead to a continuum, which is a different claim than the one the instrument is built on.
The four dichotomies are also supposed to be independent, so that any of the 16 combinations is equally plausible. In practice they are not fully independent of each other, which weakens the claim that each of the 16 types is a distinct, freestanding combination rather than a handful of overlapping tendencies.
Why do psychologists remain skeptical of the Myers-Briggs? What are the criticisms of the Myers-Briggs Type Indicator?
A review that applied a unified framework for test validity, requiring multiple independent sources of evidence rather than a single validation study, concluded that there is insufficient evidence to support the tenets and practical claims made about the instrument. That is the core of the skepticism: not that the instrument feels wrong to take, but that the case for what it measures and what it predicts has not been built to the standard modern validity research expects.
What type does not predict at work
Is Myers Briggs scientifically valid as a predictor of job performance, team fit, or career success? The evidence has generally not supported using type for hiring, placement, or team assignment decisions. Notably, the instrument’s own publisher discourages its use in hiring and screening, which is itself a statement about the limits of what type can predict.
Why an accurate-feeling description is not evidence
Myers Briggs accuracy, in the sense of feeling personally true, and predictive validity are two separate questions, and one does not answer the other.
Does the theory behind the types hold up to scientific testing?
A scientific claim earns that label by specifying, in advance, what result would prove it wrong. Judged against that standard, the theory behind the Myers-Briggs types gets a mixed verdict rather than a single grade. Some of its claims can be tested and have not held up. Others are built in a way that makes testing impossible from the start.
Is MBTI pseudoscience or just an untestable theory
The dichotomy claim is the clearest example of something testable. It predicts that people cluster into two distinct groups on each dimension, introvert or extravert, thinker or feeler, with few people sitting in the middle. That is a real prediction, and it can be checked against how traits actually distribute in the population. An analysis of the theoretical validity behind the Myers-Briggs Type Indicator found that this is exactly where the theory fails: traits show up as smooth, continuous curves rather than two separate peaks, which is the opposite of what a clean dichotomy would predict.
Type dynamics and the dominant function hierarchy sit in different territory. There is no independent way to measure which function sits highest in someone’s hierarchy apart from the four letters that already assume it. Without a separate measurement, there is no result that could disconfirm the hierarchy. That is the heart of mbti falsifiability as a problem: a claim with no possible failure condition has not been tested, it has been assumed.
