The Big Five Personality Traits: The Best-Evidenced Model in Personality Psychology, and the Engine It Doesn’t Have

Contents
  1. What the Big Five Traits Are
  2. Where the Five Came From
  3. The Argument About the Fifth Name
  4. What Travels, and What Doesn't
  5. The Trait That Predicts Things
  6. The Situation Problem
  7. What Is Inherited, and What That Word Means
  8. The Sixth Factor
  9. How to Read a Score Without Overreading It
  10. Who Uses This, and What It Costs
  11. The Objection That Was Never Answered
  12. What a Description Is For
Big Five personality traits — the OCEAN model
  • The model

    • Five continuous dimensions — openness, conscientiousness, extraversion, agreeableness, neuroticism — each breaking into narrower facets. Nothing in the model sorts people into kinds.

  • What it has earned

    • Conscientiousness rated in childhood in the 1920s tracked how long those children lived, and the trait-outcome literature came back 87 percent intact under preregistered replication.

  • What the numbers license

    • The best-replicated link — conscientiousness to job performance — sits near .22, which leaves more than ninety percent of the outcome elsewhere. It ranks groups and does not reach individuals.

  • Where it gets used

    • Hiring applies it one applicant at a time anyway. The person who pays for that is never told.

Pulled out of a dictionary in 1936, then tested for eighty years against mortality records, twin registries, and the vocabularies of a dozen languages. The five-dimension model earns nearly every claim made for it, except the one most readers came for.


In 1936, two psychologists published a 171-page monograph that consisted, in its essential part, of a list. Gordon Allport, then at Harvard, and Henry Odbert, then teaching at Dartmouth, had gone through the 1925 edition of Webster’s New International Dictionary and pulled out every word in the English language that could be used to distinguish the behavior of one person from another. They found 17,953 of them.

Then they sorted the pile into four columns. The first held what they called stable personal traits — 4,504 words for enduring dispositions, the kind of thing a person is rather than the kind of thing a person is doing this afternoon. The second column took temporary states: excited, distraught, elated. The third took words that were mostly verdicts wearing description’s clothing: worthy, insignificant, admirable. The fourth was everything else, the metaphors and physical descriptors and the words that refused to sit still.

Column one was the one they marked as psychologically useful. Everything the Big Five became came out of that column. Nearly everything the model cannot do is in the three they set aside, and the distance between those two facts is the subject worth eighty years of argument.

Dartmouth’s alumni magazine noted the publication that May with local pride, observing that two of its own had gotten around to listing 17,953 words descriptive of personality. It was not obvious at the time that anything would come of it.

What the Big Five Traits Are

The Big Five is the claim that most of the reliable variation in human personality can be summarized along five broad dimensions, each continuous, each roughly independent of the others, and each measurable well enough to predict something.

Openness to experience covers imagination, aesthetic sensitivity, intellectual curiosity, and appetite for the unfamiliar. High scorers seek out complexity and abstraction and tend toward unconventional values; low scorers prefer the concrete, the tested, and the familiar. Its name is the least settled of the five, for reasons that took the field forty years to work out.

Conscientiousness covers organization, diligence, impulse control, and follow-through. It is the trait with the strongest documented links to consequential life outcomes, which makes it the model’s best evidence for itself and the natural place to test whether the whole enterprise is measuring something real.

Extraversion covers sociability, assertiveness, talkativeness, and positive emotionality. Both poles are ordinary. The research does not treat introversion as a deficit, and the trait is better understood as reward-sensitivity and social energy than as liking or disliking people.

Agreeableness covers trust, compassion, cooperation, and deference in conflict. Low agreeableness is not a synonym for cruelty; it describes a person whose default is to contest rather than accommodate, which has costs and uses.

Neuroticism covers the tendency toward negative emotion — anxiety, irritability, vulnerability to stress, emotional volatility. Some researchers prefer to name the healthy pole and call the dimension emotional stability, which is the same axis read from the other end.

Two structural facts matter more than the definitions. First, each dimension breaks down into narrower facets: conscientiousness, in the most widely used commercial instruments, splits into components including orderliness, dutifulness, achievement-striving, and self-discipline, and these do not always move together or predict the same things. A person can be exacting about deadlines and indifferent to tidiness, and the domain score averages that difference away. Second, none of the five is a category. Scores distribute continuously, which is why the model has no types in it, no sixteen or thirty-two of anything, and no printout that tells you which kind you are. The output is five positions on five ranges, and the interesting cases are usually the ones where one facet contradicts its own domain.

The five carry a mnemonic — OCEAN, sometimes CANOE — which is the only frivolous thing about them.

Where the Five Came From

The model has an unusual origin for something that ended up in the peer-reviewed literature: it started as a linguistic bet.

The bet is called the lexical hypothesis, and Francis Galton stated a version of it in 1884, suggesting that one could gauge the number of conspicuous aspects of human character by counting the words a language has for them. The reasoning is that if a difference between people matters enough, often enough, across enough generations, some language will eventually find a word for it, and the words will still be there to count. Allport and Odbert’s list was that count performed thoroughly.

Raymond Cattell took the pile next. Through the 1940s he reduced the trait terms by clustering synonyms and running the new machinery of factor analysis over ratings, arriving eventually at a sixteen-factor model that he defended for the rest of his life. Others working the same data kept getting a smaller number. Donald Fiske found five in 1949.

Then the most consequential paper in the model’s history was published where nobody would read it. In 1961, two researchers working for the U.S. Air Force at Lackland Air Force Base in Texas — Ernest Tupes and Raymond Christal — reanalyzed Cattell’s and Fiske’s correlation matrices across eight separate samples and found the same five factors coming back every time. They named them Surgency, Agreeableness, Dependability, Emotional Stability, and Culture, and published under the title Recurrent Personality Factors Based on Trait Ratings, as Aeronautical Systems Division Technical Report 61-97. It was a military technical report. It sat there for three decades, eventually reprinted in a journal in 1992, long after the field had rediscovered its contents the hard way.

Warren Norman at Michigan noticed it in 1963, replicated the five with peer-rating data, and relabeled Dependability as Conscientiousness — the name that stuck. And then the finding was forgotten again. By 1980 the five-factor result was not part of working knowledge in personality psychology, which is a strange thing to be true of a result that had by then been independently obtained at least four times.

Lewis Goldberg found it a fifth time. Running his own lexical project through the 1970s and 1980s, feeding in different adjective sets, different samples, different rotations, he kept recovering five, and he made the recovery public and repeated it until the field could not ignore it. In 1981 he gave the factors the name they now carry, choosing “Big Five” to mark their breadth rather than their importance — each factor is broad enough to contain dozens of narrower traits, which is a claim about resolution, not about rank.

The convergence that followed is the strongest single piece of evidence the model has, and it is easy to miss because it looks like bookkeeping. Paul Costa and Robert McCrae arrived at five from an entirely different direction: not from dictionary adjectives but from questionnaire items, factor-analyzing existing personality inventories built on clinical and theoretical traditions that owed nothing to the lexical project. Two research programs, different raw materials, different assumptions, different decades, landing on approximately the same five dimensions. When two methods that share no premises converge, the convergence is information. It does not prove the five are natural kinds. It does make coincidence an expensive explanation.

Cattell never accepted it. He referred to the growing consensus as the five-factor heresy and went on defending sixteen.

The Argument About the Fifth Name

Here is where the seams show, and they should be looked at, because a model whose factors are all cleanly named would be hiding something.

Four of the five settled. The fifth never did. Tupes and Christal called it Culture, and meant something like refinement — polished, artistic, cultivated. Norman kept Culture. Goldberg’s lexical work kept returning something that looked less like refinement and more like mental horsepower, and he called it Intellect: intelligent, insightful, knowledgeable, original. Costa and McCrae’s questionnaire route produced a factor with a different center of gravity again, more about receptivity to experience than about cleverness, and they called it Openness to Experience. Gerard Saucier suggested at one point that Imagination might be the honest label for the whole territory.

These are not four names for one thing. They are four different bets about what the factor contains, and the disagreement had a consequence that ought to make anyone cautious about factor-analytic results in general. When researchers went back and examined how the studies had been built, they found that investigators who favored the Intellect reading had tended to include intelligence-related adjectives in their item sets, and investigators who favored the Openness reading had tended to include imaginative and aesthetic ones. Each camp’s factor came out looking like the camp’s preference, partly because each camp had chosen what to feed the analysis. The method cannot find what the researcher did not put in.

The field’s eventual resolution was to stop choosing. Openness and Intellect appear to be two correlated but separable aspects of one broad domain — one oriented toward perceptual and aesthetic engagement, the other toward abstract and semantic engagement — and the compound label Openness/Intellect is now standard in the research literature precisely because the compound is more honest than either half. That is a real repair. It is also a reminder that the tidy five-item list has an argument inside it that took forty years and is settled by agreement rather than by discovery.

What Travels, and What Doesn’t

If the lexical hypothesis is right, the five should show up in other languages, because other languages have also been noticing differences between people for a long time. This is the model’s most ambitious claim and its most instructive failure.

The affirmative case is substantial. Lexical and questionnaire studies across German, Dutch, Italian, Czech, Hungarian, Polish, Turkish, Hebrew, Korean, Japanese, and Filipino have recovered structures recognizably like the five, with the usual caveat that the fifth factor behaves inconsistently and agreeableness sometimes shifts its contents. In 2007, a project spanning fifty-six nations found the five sufficiently comparable across them to report and compare national profiles. McCrae and Costa had already argued, a decade earlier, that trait structure was a human universal — a solid beginning for understanding personality everywhere, in their phrase.

Then someone tested it where the assumption was weakest.

Between January 2009 and December 2010, a research team working with the Tsimane’, a society of forager-horticulturalists in the Bolivian Amazon, administered a translated forty-four-item Big Five inventory to 632 adults. The translation had been built carefully, back-translated and comprehension-tested. The results did not contain the Big Five. Internal consistency ran below standard benchmarks for most of the five factors. Fit to the model, tested by rotating the Tsimane’ structure onto a U.S. structure, produced an average congruence coefficient of 0.62, where 0.90 is the conventional threshold for saying two structures are the same. A separate sample of 430 adults rating their spouses did no better.

The most telling detail is small. A Tsimane’ respondent who described himself as reserved also tended to describe himself as talkative. Not because the answers were careless, but because the concept doing the work in the questionnaire — extraversion, a single dimension along which reserved and talkative are opposite ends — was not the concept organizing his self-description. What the analysis found instead were two broad factors the researchers named prosociality and industriousness, and those two were psychometrically better behaved among the Tsimane’ than the imported five: higher internal reliability, better response stability, independence from each other, and agreement between self-reports and spouse reports.

The obvious explanation is literacy and schooling, and the researchers checked it. Stratifying by education, by Spanish fluency, by sex, and by age changed nothing. Removing items that might have provoked response bias produced factors that gestured at extraversion, agreeableness, and conscientiousness, and the overall fit stayed poor.

There is a subtler version of the same problem, and it produces the most elegant finding in this literature. Chinese psychologists skipped translation and built an instrument from the ground up out of locally derived constructs, and standardized it on nearly two thousand Chinese adults. Factored jointly against the standard five-factor questionnaire, most of its scales lined up with the familiar domains. One factor did not: a dimension its developers called Interpersonal Relatedness, assembled from harmony, face, and the reciprocal-obligation concept of ren-qing, which loaded on none of the five. The symmetry is the part worth keeping. Openness, the fifth Western factor, was correspondingly absent from the indigenous instrument’s scales. Two instruments, each unable to see one of the other’s dimensions, each complete by its own lights. And when Interpersonal Relatedness later turned up in Singaporean, Hawaiian, and mainland American samples, the test was renamed to drop the word Chinese, which is what it looks like when a culture-specific finding turns out to have been a gap in the other side’s vocabulary.

A second study, in 2019, took the problem to scale: twenty-nine face-to-face surveys, 94,751 respondents, twenty-three low- and middle-income countries. The finding was that the standard personality questions frequently failed to measure the traits they were built to measure, with validity low enough to compromise conclusions drawn from them. Respondents endorsed items pointing in opposite directions — agreeing both that they keep things in order and that they can be somewhat careless — and the measurement error correlated with cognition, which correlates with income, which means the noise was not random with respect to the very outcomes such surveys are usually run to study.

And then the twist that most summaries of this literature leave out. When the same research group looked at internet-collected data from the same countries, the structure looked much more like the familiar five. Whatever was breaking was not simply culture. Something about the format — the face-to-face survey, the enumerator explaining an item in her own words, the interview conditions — was doing damage that the same population did not exhibit when answering on a screen. The instrument, not only the people, is part of what gets measured.

The honest summary is narrower than either camp’s slogan. Something five-ish recovers robustly in literate, schooled, questionnaire-fluent populations across many languages, which is a real and non-trivial finding. Outside those conditions the structure degrades badly, and the field cannot yet cleanly separate how much of the degradation is the people, the translation, the format, or the act of self-rating itself.

The Trait That Predicts Things

Conscientiousness is where the model stops being a description and starts being an instrument, and the evidence is old enough to have outlived the people it was collected from.

Start with the strangest finding in the literature. In the 1920s, the psychologist Lewis Terman began following a cohort of gifted California children, collecting teacher and parent ratings on each of them. Decades later, researchers went back to those childhood ratings and matched them against death records. Children who had been rated as more conscientious were still alive at ages where their less conscientious classmates were not. The effect survived controls for sex and for parental divorce, two known contributors to shortened life.

A rating written on a form by a schoolteacher in the 1920s carried information about a death date in the 1980s. Whatever the questionnaire was measuring, it was not only measuring how people like to describe themselves.

The finding has held up as the literature has grown. A 2004 meta-analysis traced the mechanism through behavior, showing that conscientiousness relates systematically to the health behaviors that dominate mortality statistics: smoking, drinking, drug use, risky driving, adherence to treatment, exercise. The pathway is unmysterious. Conscientiousness describes how reliably a person does the boring thing, and forty years of boring things include the ones that turned out to matter.

In work, the pattern is the most replicated result in personnel psychology. Three independent meta-analyses, using different primary studies and different decision rules, estimated the true-score correlation between conscientiousness and job performance at .20, .22, and .25. No other one of the five generalizes across occupational categories the way this one does. A 2019 synthesis pulled together ninety-two separate meta-analyses covering 175 distinct occupational variables, more than 2,500 studies, and over 1.1 million participants, and found conscientiousness pointing in the desirable direction for 98% of them.

In education, a 2009 meta-analysis found conscientiousness predicting academic performance at a magnitude comparable to measured intelligence, and at the university level slightly exceeding it for grade point average. That result deserves a moment, because it inverts the folk model of school. Past the admissions gate, the trait that describes whether you do the reading competes with the trait that describes how fast you understand it.

And now the part that most popular accounts skip, because it is the part that keeps the rest honest.

A correlation of .22 is a real relationship and a small one. Squared, it accounts for under five percent of the variance in the outcome. What it means in practice is that if you sort a thousand employees by conscientiousness score and look at their performance ratings, the high group will visibly outperform the low group as a group, while any individual you point at may sit anywhere. It supports statements about populations. It does not support statements about the person in front of you, and the entire commercial apparatus built on personality testing consists, in large part, of selling the second thing using evidence for the first.

The field also stress-tested its own outcome literature, which is more than most subfields have done. In 2019 a preregistered project attempted high-powered replications of seventy-eight previously published associations between the five traits and life outcomes, drawn from a comprehensive review, with a median sample of 1,504 and more than 6,100 adults in total. Eighty-seven percent came back statistically significant in the expected direction. The replicated effects were typically 77% as strong as originally reported. The author’s own reading was appropriately unsentimental: the literature holds up better than several neighboring fields, and the shortfall from perfect replication is still consistent with some false positives, selective reporting, and publication bias in the original record.

That is the shape of the honest claim. The associations are real, they mostly replicate, they are smaller than the first papers said, and they are more useful for describing a thousand people than for describing one.

The Situation Problem

There is an objection every reader has already thought of, and it once nearly destroyed this entire field.

In 1968 Walter Mischel published Personality and Assessment, reviewed the evidence linking trait scores to actual behavior, and reported that the correlations rarely exceeded .20 to .30. He gave the number a name that stuck for a generation, the personality coefficient, and drew the conclusion the number seemed to support: if knowing someone’s trait score barely improves your guess about what they will do, then broad dispositions are not the right unit, and what actually determines behavior is the situation. A talkative person is quiet at funerals. A conscientious person misses a deadline in a week from hell. Behavior looked too inconsistent across contexts to belong to persons at all.

The critique was devastating and the field spent two decades absorbing it. What eventually resolved it was not a defense of traits but a correction of the arithmetic.

Seymour Epstein pointed out in 1979 that the comparison had been unfair in a specific, technical way. A single behavior on a single occasion is a one-item measurement of anything, and one-item measurements are unreliable by construction. Nobody would estimate a baseball player’s ability from one at-bat, then conclude that batting ability does not exist because the correlation between one at-bat and another is low. Aggregate the behavior across occasions and the correlations rise steeply, because the noise cancels and the tendency accumulates. Traits do not predict what a person will do at 3 p.m. on Thursday. They predict what a person will have done by December, which is a different question and the one that matters for almost everything anyone cares about.

The most useful reframing came from William Fleeson in 2001, and it changes what a trait score means. Sample people’s behavior many times a day for weeks and you find that the average person expresses, at some point, nearly every level of every trait: everyone is talkative sometimes and withdrawn sometimes, exacting sometimes and slapdash sometimes. Within-person variability in trait expression turns out to be large, often larger than the differences between people. What distinguishes individuals is not a fixed setting but the shape of their distribution — where its center sits, how wide it spreads, and which situations pull them where. And the center of that distribution is, in Fleeson’s phrase, among the most predictable variables in psychology.

So a conscientiousness score is not a description of how you behave. It is an estimate of the middle of a range you move through constantly. This is a more modest claim than the results page implies and a considerably more defensible one, and it dissolves the objection without pretending the objection was wrong. Mischel was right that behavior is wildly inconsistent. He was wrong that the inconsistency leaves nothing stable to measure.

What Is Inherited, and What That Word Means

Twin studies have been asking a version of one question for fifty years: identical twins share all of their DNA, fraternal twins share about half, and if a trait is heritable, the identical pairs should resemble each other more. Applied to the five, this design has produced a consistent answer. Roughly 40 to 60 percent of the variance in each of the five traits is attributable to genetic variation. The estimate has held across countries, cohorts, and instruments.

The second finding from the same designs is the one that unsettled people. The remaining variance is almost entirely non-shared. Growing up in the same household, with the same parents, the same rules, the same books, and the same dinner table, does not make siblings resemble each other in personality to any degree these studies can reliably detect. Shared environment sits near zero. Whatever the other half of personality is, it is not the family in the aggregate; it is something more local and more particular, and after five decades the field still describes it mostly by what it is not.

Then molecular genetics arrived and made the picture harder. If personality is 40 to 60 percent heritable, the specific genetic variants should be findable. They have been much harder to find than expected. A 2015 study using genome-wide data from 5,011 adults and over half a million genetic markers could account for about 15% of the variance in neuroticism and 21% in openness, and found no statistically significant heritability at all for extraversion, agreeableness, or conscientiousness. Later and much larger studies have identified variants associated with these traits, particularly neuroticism, but the gap between what twin designs imply and what the genome has yielded remains substantial and is not yet closed.

Several candidate explanations are respectable, and the field holds them simultaneously. Common-variant methods systematically miss rare and structural variation, and typically recover roughly half of a twin estimate for that reason alone. Personality may be extraordinarily polygenic, spread across thousands of variants each too small to detect at current sample sizes. And twin designs may themselves inflate the number, if non-additive genetic effects or subtle assumption violations are doing work that the model assigns to heritability.

One clarification is worth more than all of the numbers, because the numbers are so reliably misread. Heritability is a statement about a population, not a person. To say conscientiousness is 45% heritable does not mean 45% of your conscientiousness came from your parents and the rest from your upbringing. It means that in the population studied, under the conditions that prevailed there, about 45% of the differences among people could be statistically attributed to genetic differences among them. Change the population or the conditions and the fraction changes. A trait can be highly heritable and highly malleable at the same time, which is not a paradox and is one of the more useful things this literature has to teach.

One further finding sits awkwardly beside every number in this section. The traits move. Across adulthood, and on average, conscientiousness and agreeableness rise, neuroticism declines, and openness drifts down in later decades. People become, in the aggregate, somewhat more reliable and less easily distressed as they age. Whatever the genome is doing, it is not holding these numbers still.

The Sixth Factor

The most serious structural challenge to the five did not come from critics of the trait approach. It came from researchers doing the lexical work more thoroughly, in more languages, and finding one more thing.

In 1999, a lexical study of Korean personality vocabulary recovered variants of the usual five and also something else: a clearly interpretable additional factor contrasting words like truthful and frank against words like sly and cunning. Anomalous sixth factors had surfaced in earlier non-English studies and been treated as noise, one at a time, by researchers who had no reason to connect them. The Korean result was clean enough to make the pattern visible in retrospect.

Michael Ashton and Kibeom Lee went back and reanalyzed eight lexical studies across seven languages — Dutch, French, German, Hungarian, Italian, Korean, and Polish — and the sixth factor came back consistently. They named it Honesty-Humility: sincerity, fairness, modesty, and the absence of greed, against boastfulness, hypocrisy, and manipulation. The resulting six-factor model is called HEXACO, and it rearranges the furniture rather than adding a chair. Agreeableness is redefined so that anger and temper sit at its low end instead of inside neuroticism, and neuroticism itself is recast as Emotionality with a different center of gravity.

The empirical argument for the sixth factor is that it does work the five cannot. Honesty-Humility predicts exploitation, unethical decision-making, counterproductive behavior at work, materialism, and power-seeking better than any single one of the five, and it accounts for the antisocial traits that the five-factor framework has always handled awkwardly.

The counterargument is specific and has not been dismissed. Honesty-Humility correlates substantially with the agreeableness domain as measured by the standard questionnaire instruments, particularly with two of its facets, straightforwardness and modesty. On that reading the sixth factor is not new territory but a piece of existing territory promoted to its own continent. The six-factor side answers with measurement: realign those facets into a separate factor and prediction improves for criteria involving deceit without hostility, by margins that were measured rather than asserted.

The dispute is live. In 2020 a leading journal ran the six-factor case as a target article with commentaries attached, and the commentaries disagreed with each other about the number of factors, about whether the lexical approach can settle such a question at all, and about whether the count matters as much as the argument implies. That is what an unresolved empirical question looks like when the field is behaving well. The five remains the default in most research, the six is standard in a substantial and growing minority of it, and anyone telling you the number of basic personality dimensions is a settled matter is describing a consensus that does not exist.

How to Read a Score Without Overreading It

Most people meet this model as five numbers on a results page, so it is worth being concrete about what those numbers can bear.

They are percentiles against a comparison group, not quantities. A conscientiousness score in the 70th percentile means roughly that seven of ten people in whatever sample the test was normed against scored lower. It does not mean you possess 70 units of anything, and it is not comparable across instruments normed on different populations. Change the reference group and the number moves without anything about you changing.

The middle is not a finding. Scores near the fiftieth percentile are the most common result and the least informative one. They can mean genuine moderation, or a mix of high and low facets averaging out, or an instrument too short to resolve the question. Only the tails carry much signal, which is inconvenient given how many people land in the middle and go looking for meaning there.

Facets beat domains for anything specific. Two people with identical conscientiousness scores can differ completely underneath, one high on orderliness and low on achievement-striving, the other the reverse. Where a question is specific, the facet is the level with information in it, and any instrument reporting only five numbers has discarded that level. The five-number summary is a compression, and compressions lose what they were built to lose.

The instrument matters, and the good ones are not always the expensive ones. Public-domain item pools developed in the research community are the basis of many free tests and are, item for item, as well validated as commercial inventories. The relevant question about any test is not its price but whether it reports facets, states its norming sample, and reports its reliability. Most consumer versions do none of the three.

Self-report has known, directional biases. People answer as the role they occupy, as the person they are trying to become, and as the season they are in. A depressive episode inflates neuroticism. A demanding job inflates conscientiousness. Observer ratings from people who know you correlate respectably with self-reports but not perfectly, and where they diverge, neither is automatically the truth.

And your score will move. Test-retest stability over short intervals is high, over decades considerably lower, partly from measurement noise and partly because the underlying trait genuinely changed. A number that moved is not necessarily a number that was wrong.

Nothing here should be read as a verdict on any individual. Trait descriptions are statements about distributions of people, and a person who recognizes themselves in one is still the only available expert on the ways the description fails them.

Who Uses This, and What It Costs

The Big Five did not stay in the journals. It is the backbone of the personality-testing industry in hiring, and its language has spread into insurance modeling, marketing segmentation, and any commercial context where a large population needs sorting cheaply.

The validity evidence supports part of this and not the part that matters most to the person being sorted. A conscientiousness correlation in the low twenties genuinely improves selection across a thousand hires. It says almost nothing reliable about any single applicant, and a hiring decision is always a single applicant. This is not a subtle statistical caveat; it is the central fact, and the industry’s business model requires it to stay in the footnotes. Personnel psychologists have said so themselves: a prominent 2007 paper by a group of the field’s own senior figures reconsidered the use of personality tests in selection and concluded, from inside the discipline, that the observed validities are lower than practitioners believe and that the tests are being oversold.

There is a second problem the research has documented and the market has not solved. When a questionnaire’s result affects whether someone eats, people manage their answers. Impression management under high stakes is measurable, it is not evenly distributed across candidates, and the applicants most likely to answer honestly are the ones least coached about what the instrument rewards.

So the cost lands somewhere specific. Someone did not get the job. There was no interview to argue in, no claim to rebut, nothing to appeal, because a profile arrived before she did and the profile was assembled from evidence that was never valid at the level of one person. She will not learn this happened. The literature that supports the practice contains, in its own numbers, the grounds for saying the practice exceeds it, and that has been true for thirty years without noticeably slowing anything down. The research community mostly agrees this is a misuse. The misuse is what most people will meet the model as.

The Objection That Was Never Answered

In 1995, Jack Block published a paper in Psychological Bulletin arguing that the five-factor approach had been granted a theoretical status it had never earned. Block was a distinguished personality psychologist at Berkeley, worked from inside the field, called himself a contrarian, and made his case in detail across nearly thirty pages.

The core of it: factor analysis does not discover the structure of anything. It summarizes covariation among the variables it is given, and the five factors are therefore a summary of how people use a particular set of personality-descriptive words. Feed the analysis a different set and a different structure comes out, which is exactly what the fifth-factor argument had already demonstrated. The model makes no prediction about why these five should exist rather than four or seven, offers no account of what produces them, describes no developmental process, and specifies no mechanism. It is a taxonomy of language that has been repeatedly mistaken for a map of the mind. Block’s broader complaint, which is usually forgotten in favor of the technical parts, was that personality psychology had made variables its object of study, and lost the people.

The replies came in the same issue. Costa and McCrae answered. Goldberg and Saucier answered under a title that concedes more than it means to: So what do you propose we use instead?

That question is a good one, and it has never been answered either, and both facts are load-bearing. Block was substantially right that the model is atheoretical in origin, and its principal developers have largely conceded the point while arguing that a useful taxonomy does not require a prior theory — that chemistry had a periodic table before it had an account of why the periods repeat. The defense is pragmatic and it is strong: the five describe the space efficiently, predict outcomes reliably, and have let thousands of researchers accumulate findings that add up, which no competing framework has matched. Block returned to the argument in 2001 and again in 2010. He did not persuade the field. He was also never refuted.

Serious work has gone into supplying the missing engine since. Colin DeYoung has proposed that the five sit under two higher-order metatraits, Stability and Plasticity, with candidate neurobiological substrates, and has built out a theory that treats the traits as parameters of a goal-directed system instead of summary labels. It is the most credible attempt at what Block said was missing, it is one theory among several, and it has not become consensus.

So the situation, thirty years on, is that the best-validated descriptive model in personality psychology still cannot tell you where its own five dimensions come from. The measurement outran the understanding and has stayed ahead of it.

What a Description Is For

The Big Five is the most successful thing personality psychology has built. It replicates across many languages, predicts mortality and job performance and academic achievement at effect sizes that survive preregistered replication, shows substantial heritability, and has supported a cumulative research literature for four decades. Any account that treats it as pseudoscience is not describing the evidence.

It is also, and permanently so far, a description with no engine inside it. It tells you where a person sits on five ranges. It does not tell you why they sit there, what it is like to sit there, or what to do about it. Those three questions are the reasons most people go looking in the first place, and the model’s answer to all three is silence — not a modest answer, not a partial one, silence.

The gap is worth sitting with, because the resolution on offer is false. There is no version of the research program in which enough further measurement turns into an explanation. Better instruments produce better-calibrated numbers, and a well-calibrated number is not an account of a life. Someone who arrives with the question why am I like this and leaves with five percentiles has been answered accurately and has not been answered. The accuracy does not reduce the gap. It is what the gap is made of.

That gap explains something the research literature finds embarrassing. The four-letter type languages, which measure worse and predict less, keep the audience anyway — because they operate in the register the questions were asked in, offering a felt account of how a mind works where the trait model offers a position on a range. That is a description of a genuine trade, not a recommendation: the better-evidenced framework will not tell you why, and the more satisfying one will tell you why using machinery whose evidentiary status is much weaker. Knowing which of the two failures you are buying is most of the skill, and the companion guide to the sixteen-type framework reaches the same conclusion from the other direction.

Which returns to the two men and the dictionary.

Allport and Odbert sorted 17,953 words into four columns and marked one as useful. The Big Five is what eighty years of rigorous work extracted from column one, and it is an extraordinary yield: 4,504 words for stable dispositions, compressed to five dimensions, that turned out to carry information about death records and job performance and twin resemblance. Nobody in 1936 could have promised that.

The other three columns were not empty. Column two held the temporary states, which is where most of a life is actually lived. Column three held the evaluations, the words that judge rather than describe, which is what people mostly do to each other and what any honest account of personality has to reckon with instead of discarding. Column four held the metaphors, and metaphor is how one person has always managed to convey to another what being themselves is like.

The model works because those columns were set aside. The questions it cannot answer are still in them.

This article is available at https://cinemawords.com/en/big-five-personality-traits/

Continue Reading

Advertisement