HEXACO: The Six-Factor Personality Model
The HEXACO model is a six-factor lexical personality taxonomy developed by Kibeom Lee and Michael C. Ashton in which Honesty-Humility is added as a factor distinct from a redefined Agreeableness, Emotionality replaces Big Five Neuroticism with revised content, and Extraversion, Conscientiousness, and Openness retain forms close to their Big Five counterparts. The model's strongest empirical case is cross-language lexical recurrence of a Honesty-Humility-like factor plus incremental prediction of integrity-relevant outcomes — fairness, sincerity, modesty, exploitation, workplace deviance, and Dark Triad-adjacent variance — that broad Big Five Agreeableness blurs. Whether HEXACO should displace the Big Five as the default trait taxonomy, including for AI-personalization and alignment-adjacent applications, remains scope-dependent and contested.
Coverage note: verified through May 2026.
1. The model and its six factors
HEXACO emerged from a series of lexical studies in the 1990s and early 2000s in which trait-descriptive adjectives from multiple natural languages were factor-analyzed independently of the English-language work that produced the Big Five. The decisive comparative analysis is Ashton, Lee, Perugini, Szarota, de Vries, Di Blas, Boies, and De Raad (2004) — a joint examination of psycholexical studies across seven languages — which reported that a six-factor solution recurred more cleanly across languages than the canonical five-factor solution, and that the additional sixth factor consistently loaded on adjectives related to sincerity, fairness, lack of greed, and modesty.
The model is named HEXACO for the initial letters of its six domains: Honesty-Humility, Emotionality, eXtraversion, Agreeableness (versus anger), Conscientiousness, and Openness to Experience. Each domain is divided into four facets in the standard scoring; an additional interstitial facet, Altruism (versus Antagonism), loads on H, E, and A and is reported separately. The current facet structure was finalized in Lee and Ashton (2006), which added the Altruism interstitial facet and rebranded the instrument as the HEXACO-PI-R.
The six domains, with the standard facets used in HEXACO-PI-R scoring, are summarized below.
| Domain | Definition (positive pole) | Facets | Closest Big Five mapping |
|---|---|---|---|
| Honesty-Humility (H) | Tendency to be fair and genuine in dealing with others; to avoid exploiting them for personal gain | Sincerity, Fairness, Greed-Avoidance, Modesty | Partial overlap with NEO-PI-R Agreeableness facets of Straightforwardness and Modesty; partial overlap with Conscientiousness facet of Dutifulness; no clean single Big Five home |
| Emotionality (E) | Tendency to experience fear of physical danger, anxiety, dependence on social support, sentimental attachment | Fearfulness, Anxiety, Dependence, Sentimentality | Overlaps with Big Five Neuroticism but with anger/irritability content removed (relocated to low HEXACO Agreeableness) |
| eXtraversion (X) | Tendency to feel positive about oneself, to lead social groups, to enjoy social interaction and lively activity | Social Self-Esteem, Social Boldness, Sociability, Liveliness | Close to Big Five Extraversion |
| Agreeableness (A) | Tendency to forgive, to be lenient in judging others, to cooperate even when exploited, to control temper | Forgiveness, Gentleness, Flexibility, Patience | Overlaps with Big Five Agreeableness but centered on patience and low anger rather than on soft-heartedness or trust |
| Conscientiousness (C) | Tendency to organize time and surroundings, to work toward goals, to deliberate carefully, to strive for accuracy | Organization, Diligence, Perfectionism, Prudence | Close to Big Five Conscientiousness |
| Openness to Experience (O) | Tendency to be absorbed in beauty, to be inquisitive about diverse domains, to be imaginative, to be interested in unusual ideas | Aesthetic Appreciation, Inquisitiveness, Creativity, Unconventionality | Close to Big Five Openness/Intellect, typically narrower on abstract intellect |
Two features of the structure are easy to lose in summaries. First, the rotation that produces Honesty-Humility also redefines Agreeableness: the hostile/irritable content that lives in low Big Five Agreeableness is reallocated, with anger tracking low HEXACO Agreeableness and exploitation/greed/manipulation tracking low HEXACO Honesty-Humility. Second, HEXACO Emotionality is not Big Five Neuroticism with a friendlier label — it removes anger and adds sentimentality and dependence, producing a domain about fear, attachment, and emotional vulnerability rather than negative emotionality in general. The result is a taxonomy whose six factors do not map one-to-one onto five Big Five domains; the differences extend beyond simply "adding H."
A semantic note that matters for any careful reader: Greed-Avoidance is a positively keyed facet of Honesty-Humility. High Greed-Avoidance describes someone uninterested in luxury, social status, and wealth-based prestige; low Greed-Avoidance describes someone motivated by material gain and status signaling. Phrases like "low greed-avoidance" therefore denote more greed, not less, and the surrounding literature is occasionally loose about this. The same applies to the other H-facet labels: high Modesty means absence of conceit, high Sincerity means absence of manipulation, high Fairness means absence of fraud — all keyed so that higher scores indicate the integrity-relevant pole.
For broader context on personality taxonomy, see Big Five Personality Model, Psychometrics, Construct Validity in Psychometrics, and Lexical Hypothesis in Personality Psychology.
2. Lexical and cross-cultural evidence
The case for a sixth factor rests on what the lexical hypothesis can and cannot establish. The lexical hypothesis holds that important individual differences become encoded as single trait-descriptive words in natural language, and that factor-analyzing such words should recover the trait structure that matters most in social life. Five-factor researchers from Goldberg onward took this seriously, but the original English-language studies privileged English adjective pools selected from earlier English-language work. The HEXACO program asked whether factor-analyzing the native trait lexicon in each language would converge on five factors, six factors, or something else.
Ashton, Lee, and Goldberg (2004), working with a comprehensive English adjective set, found that a six-factor solution included a recognizable Honesty-Humility-like factor with sincerity, fairness, and modesty content that did not load cleanly on any of the five canonical factors. The cross-language pattern in Ashton et al. (2004) is more decisive: across the languages reanalyzed jointly, the sixth factor showed reasonable recovery, while a five-factor solution forced sincerity-fairness-modesty content to be distributed across Agreeableness, Conscientiousness, and a few other domains in inconsistent ways. A subsequent review by Ashton and Lee (2007) consolidates the conceptual case, arguing that the six-factor solution can be theoretically grounded in two evolutionary pressures: reciprocal altruism (mapping to H and A) and kin altruism (mapping to E).
Two honest qualifications belong on any serious treatment. First, lexical studies are not ontological proofs. Factor count is sensitive to adjective selection, the inclusion of evaluative versus descriptive terms, the rotation chosen, the extraction method, and the linguistic structure of the source language. The same adjective space can yield three-, five-, six-, or seven-factor solutions depending on these choices; the question is which solution is most consistent across decisions, not whether the data uniquely picks one number. The HEXACO claim is that six factors are more consistent across language, rotation, and sample than five — a defensible empirical claim, not a discovery of nature's joints.
Second, the cross-cultural case has been stronger in some language families than in others. De Raad et al. (2010) argued in a 14-language reanalysis that even three-factor solutions show more cross-language consistency than five or six factors do, on the methodological ground that broader factors are more robust under translation. Ashton and Lee (and others) have responded that broader factors are trivially more robust because they combine more variance, and that the substantive question is whether the additional H factor recovers a coherent and replicable content cluster — which they argue it does. The debate is unresolved in the strong sense, but the asymmetry of evidence is real: studies that test for a Honesty-Humility-like factor regularly find one; studies that do not test for it do not.
The defensible summary is narrower than a victory claim. Lexical work across multiple languages supports a recurring sixth factor anchored in sincerity, fairness, greed-avoidance, and modesty content, and HEXACO is the most fully developed taxonomy organized around that finding. It does not establish that nature has exactly six factors. It establishes that a six-factor rotation separating H from Agreeableness produces a more interpretable and cross-linguistically stable structure than the canonical five-factor solution.
3. Measurement instruments
The measurement family is now broad enough that articles regularly conflate instruments. They are not interchangeable, and they have different publication histories, item counts, intended uses, and reliability properties.
| Instrument | Items | Facets | Primary source | Intended use |
|---|---|---|---|---|
| HEXACO-PI | 192 | 24 (no Altruism) | Lee & Ashton (2004) | Original full instrument; now superseded by the PI-R |
| HEXACO-PI-R (200) | 200 | 25 (24 + interstitial Altruism) | Lee & Ashton (2006) (revision adding Altruism facet and observer form) | Research-grade facet-level assessment when length is acceptable |
| HEXACO-100 | 100 | 25 | Lee & Ashton (2018) (psychometric properties paper) | Default research-grade self-report; standard for most large-scale studies |
| HEXACO-60 | 60 | Six domains only (facet scales too short for reliable facet-level analysis) | Ashton & Lee (2009) | Short form for time-constrained settings; domain scores only |
| Observer-report HEXACO-PI-R | 200 | 25 | Lee & Ashton (2006) (observer form) | Informant assessment paired with self-report |
Official scoring keys and item-level materials are maintained at hexaco.org.
Two practical points matter for anyone choosing an instrument. The HEXACO-60 is a screening tool, not a facet-level instrument: its facet scales are too short to support reliable inferences about within-domain structure, and the 2009 introduction explicitly recommends it for situations where only domain scores are needed. The HEXACO-100 is now the default research instrument because it preserves full facet structure (4 items per facet × 25 facets) while remaining short enough for most survey contexts; the 200-item PI-R is reserved for studies that need maximum facet-level internal-consistency reliability or that are validating new facet-level hypotheses. Lee and Ashton (2018) report internal-consistency reliabilities for HEXACO-100 facet scales typically in the .70-.80 range and domain reliabilities in the .80-.90 range, comparable to other widely used personality inventories.
Anyone working with HEXACO should also be aware that the instruments are scored with a 5-point Likert response format (strongly disagree to strongly agree), that some facet content was modified between the original PI and the PI-R (most notably the addition of the Altruism interstitial facet in 2006), and that the observer-report form uses third-person rephrasing of the self-report items. None of this is exotic, but the differences matter for cross-study comparison and for any attempt to merge HEXACO results across data collected with different forms. Treating "the HEXACO" as a single instrument across the 2004–2018 literature is a recurring source of error.
For methodological background, see Item Response Theory, Computerized Adaptive Testing, and Common Method Variance.
4. Predictive validity and the Honesty-Humility case
The strongest part of the HEXACO empirical record is the incremental predictive validity of Honesty-Humility over Big Five measures for integrity-relevant outcomes. The general pattern is consistent: across workplace deviance, counterproductive work behavior, integrity-test scores, unethical decision-making, fraud, theft, exploitation in economic games, and Dark Triad-adjacent traits, Honesty-Humility predicts as well as or better than any single Big Five domain, and typically adds meaningful incremental validity beyond the Big Five even when Big Five scores are entered first.
Early demonstrations include Ashton and Lee (2008) on social-attitude variables (right-wing authoritarianism, social dominance orientation, materialism); Marcus, Lee, and Ashton (2007) on integrity-test convergence; and Lee, Ashton, and de Vries (2005) on workplace delinquency. The accumulated evidence is summarized by Pletzer, Bentvelzen, Oostrom, and de Vries (2019), a meta-analysis of HEXACO domains and counterproductive work behavior reporting Honesty-Humility as the strongest predictor of all six domains, with effect sizes that substantially exceed those for Conscientiousness or Big Five Agreeableness in comparable analyses. Lee, Berry, and Gonzalez-Mulé (2019) extends the predictive record to job-performance-adjacent criteria in the organizational literature.
The careful reading of this evidence requires distinguishing three nested questions.
Does H predict integrity outcomes? Clearly yes, across a wide outcome family and across self-report, observer-report, and behavioral criteria.
Does H predict integrity outcomes more strongly than any single Big Five domain? Yes for most criteria, with the caveat that Big Five Conscientiousness sometimes matches or approaches H on broadly defined work-related outcomes that include reliability and rule-following.
Does H add incremental validity over a facet-level Big Five baseline? This is the question most relevant to the taxonomy debate, and it is the one least often asked directly. NEO-PI-R Agreeableness contains a Straightforwardness facet and a Modesty facet that overlap conceptually with HEXACO H facets, and NEO-PI-R Conscientiousness contains a Dutifulness facet that overlaps with low Fairness. The available evidence is mixed but tilts toward genuine incremental validity: H correlations with workplace deviance, unethical behavior, and Dark Triad measures typically remain meaningful after controlling for the relevant Big Five facets combined. That is not the same as showing that no Big Five facet combination can match H — it is logically possible that a weighted combination of NEO facets approximates the H content — but it establishes that H is not redundant with any single existing Big Five facet, and that the relevant variance is not captured by broad Big Five domain scores.
A serious counter-argument is that several apparent H advantages may be inflated by criterion contamination. Many integrity criteria are themselves self-report measures with item content that overlaps semantically with H predictor items. A self-report scale asking whether the respondent has ever stolen from an employer is closer in surface content to "I would be tempted to use counterfeit money" (a Greed-Avoidance item) than to "I see myself as someone who tends to find fault with others" (a Big Five Agreeableness item). Where studies use behavioral criteria — economic-game defection, observed task cheating, archival workplace records, informant reports of misconduct — the H advantage typically survives but with smaller effect sizes than self-report-to-self-report studies suggest. The Pletzer meta-analytic work partly addresses this by reporting effects separately for self-report and other-report criteria, but the broader field still under-uses behavioral and archival outcome measures.
The defensible summary is that Honesty-Humility provides clear incremental predictive value over Big Five domain scores for integrity-relevant outcomes, and probably retains incremental value over relevant Big Five facets — though the latter case is less unambiguously established and is sensitive to method variance. For the practitioner asking "should I add H if I am already measuring the Big Five and I care about exploitation, fairness, or integrity?", the answer is straightforwardly yes. For the theorist asking "does H demonstrate that the Big Five is wrong?", the answer is no — it demonstrates that broad Big Five domain scores blur a useful axis.
Related concepts: Integrity Testing, Counterproductive Work Behavior, Workplace Deviance, Behavioral Economics Games.
5. Big Five versus HEXACO: a scope-dependent debate
The five-versus-six debate is often posed as a winner-take-all question — which taxonomy is the correct one? — but the more useful framing is to ask what each taxonomy is good for, where they diverge, and what the consequences of each choice are.
Several considerations matter.
Coverage and parsimony. Both taxonomies are broad-bandwidth, low-fidelity models of personality. Big Five with five domains and (in NEO-PI-R) thirty facets aims to cover most of the trait-descriptive variance in adjective space; HEXACO with six domains and twenty-five facets aims at the same coverage with a different rotation. The five-factor solution is more parsimonious by one factor; the six-factor solution is more parsimonious in its treatment of the sincerity-fairness-greed-modesty cluster, which under Big Five gets distributed across Agreeableness and Conscientiousness in ways that depend on the instrument.
Instrument infrastructure. The Big Five has a much larger installed base. Decades of NEO-PI-R, BFI-2, IPIP, and TIPI work, plus accumulated norms, longitudinal datasets, and cross-cultural validation studies, give Big Five-based research enormous comparative leverage that HEXACO does not match. This is not evidence of structural superiority, but it is a real cost for switching to HEXACO when interoperability with existing data matters.
Predictive validity profile. As above, HEXACO Honesty-Humility outperforms Big Five domains for integrity-relevant outcomes, sometimes substantially. For other outcome families — subjective well-being, academic achievement, broad job performance, relationship quality — the two taxonomies produce comparable predictive validity, with the relevant variance loading on Conscientiousness, Extraversion, Emotionality/Neuroticism, or Openness in similar ways. The HEXACO advantage is concentrated; it is not uniform across criteria.
Cross-cultural recovery. Lexical evidence favors six factors on average across the language samples used in the original HEXACO program, with caveats noted in §2. Big Five recovery is also reasonable cross-culturally but treats the H cluster inconsistently across instruments and translations.
Theoretical interpretability. HEXACO researchers, especially in Ashton and Lee (2007), argue that the six factors map onto identifiable evolutionary pressures — H and A as components of reciprocal altruism, E as kin altruism, X/C/O as separate dimensions related to engagement with social, task-related, and intellectual environments. Big Five researchers offer their own theoretical frameworks (e.g., DeYoung's metatrait/aspect work). Neither is a knock-down argument; both are post-hoc theoretical organizations.
| Decision context | Recommended taxonomy | Reason |
|---|---|---|
| Selection or screening for integrity-sensitive roles | HEXACO (or H added to Big Five) | H captures distinctive integrity-relevant variance |
| Cross-study comparison with the broad personality literature | Big Five | Installed base is larger |
| Cross-cultural population studies | HEXACO | Better lexical recovery on average |
| Research on Dark Triad, exploitation, manipulation | HEXACO | Direct H axis is most interpretable |
| General-purpose user modeling without integrity focus | Either; Big Five if interop matters | No clear HEXACO advantage |
| Research where facet-level Big Five baselines are needed | Either, with caution | HEXACO H may still incrementally predict, but the gain shrinks |
The article-level conclusion is that the Big Five and HEXACO are partly competitive, partly complementary. They share most variance and most predictive validity for most outcomes; they diverge meaningfully on integrity-relevant criteria, on the treatment of exploitation/fairness/modesty content, and on the cross-cultural lexical evidence. The honest answer to "which taxonomy is correct?" is that the question is malformed: trait taxonomies are rotation choices over a high-dimensional space, and the better question is which rotation makes the variance you care about more interpretable.
For complementary framing, see Big Five Personality Model, Personality Trait Stability, and Lexical Hypothesis in Personality Psychology.
6. HEXACO and the Dark Triad
The Dark Triad — Machiavellianism, narcissism, and psychopathy — is one of the most-studied areas of personality in which HEXACO's Honesty-Humility offers clearer measurement than the Big Five does. The basic finding, demonstrated by Lee and Ashton (2005) and replicated by many subsequent studies, is that low Honesty-Humility correlates substantially with each of the Dark Triad measures, typically in the r ≈ -.40 to -.55 range with narcissism somewhat weaker than Machiavellianism or psychopathy.
The meta-analytic case is summarized by Muris, Merckelbach, Otgaar, and Meijer (2017), which integrates evidence from many studies of Dark Triad correlates and discusses HEXACO H as the personality factor most consistently associated with the antisocial core common to the three constructs. Howard and Van Zandt (2020) performed a meta-analytic comparison of HEXACO and Dark Triad measures and reported that low H explains a large fraction of the variance shared across the Dark Triad, with smaller contributions from low HEXACO Agreeableness, low HEXACO Conscientiousness, and high HEXACO Emotionality on subscale-specific patterns.
The substantive interpretation is that the Dark Triad measures, despite their disparate origins (Machiavellianism from political-philosophical work, narcissism from clinical and personality literature, psychopathy from forensic and clinical research), converge on a common antagonistic-exploitative core that HEXACO measures cleanly as low Honesty-Humility. This is useful both for parsimony — researchers studying Dark Triad outcomes can often replace three correlated measures with one — and for theory: the convergence suggests that the Dark Triad's "darkness" is largely the same construct that HEXACO researchers identified through entirely different methodological roots.
The boundary conditions matter. Low Honesty-Humility is not the Dark Triad; the equivalence is approximate. Each construct retains distinct variance that H does not capture.
- Machiavellianism includes a strategic-manipulative-cynical worldview component beyond simple antagonism.
- Psychopathy includes impulsivity, callous-unemotional traits, and reduced fear that link to HEXACO Emotionality (low fear, low anxiety) and Conscientiousness (low Prudence) in addition to low H.
- Narcissism is heterogeneous: grandiose narcissism tracks low H reasonably well, but vulnerable narcissism shows a different profile with elevated Emotionality and a weaker H link.
- The Dark Tetrad, which adds sadism to the original three constructs, shares the low-H core but also picks up specific patterns (e.g., low HEXACO Agreeableness, low Emotionality) that distinguish it from pure low-H.
The practical recommendation is that HEXACO H is the single best parsimonious measure of Dark Triad-shared variance, and that researchers interested in the specific contributions of Machiavellianism, narcissism, or psychopathy beyond their shared core should retain the specific measures. Substituting H for the Dark Triad is appropriate when the research question concerns the shared antagonistic-exploitative dimension; it is inappropriate when the question requires distinguishing the three constructs.
Related concepts: Dark Triad Personality, Antisocial Personality Traits, Machiavellianism in Behavioral Research, Sadism and Cruelty.
7. AI-adjacent applications
Whether HEXACO provides useful structure for AI-relevant work is a more speculative question than its established place in human personality research. The relevant uses fall into three families, with very different evidential bases.
7.1 User modeling for AI personalization
HEXACO could in principle be used to characterize the human users of AI systems — to predict which users are most susceptible to sycophantic outputs, most resistant to corrective challenges, most vulnerable to manipulative persuasion attempts, or most likely to over-rely on AI advice. The conceptual case is straightforward: H is associated with resistance to bribery and exploitation, low H with susceptibility to flattery and material-incentive appeals, and these mechanisms plausibly transfer to AI-mediated interactions. The empirical case is much thinner. Most work on personality and AI use relies on Big Five measures rather than HEXACO; the Psychometric Correlates of AI Interaction Styles entry surveys what is known and how indirect most of it is.
The strongest current evidence for HEXACO specifically in this area is Liang et al. (2025), which found that HEXACO Honesty-Humility, Agreeableness, and Conscientiousness all negatively predicted generative-AI academic misconduct, with personality explaining more variance than attitudes did once both were in the model. That is consistent with H being relevant to AI-mediated ethical behavior, but it does not establish that HEXACO outperforms Big Five for AI personalization across the broader range of relevant outcomes. The question of whether HEXACO features beat a Big Five baseline for prediction of AI-interaction outcomes has not been answered in a preregistered, multi-criterion study to date.
7.2 Persuasion resistance and deception susceptibility
Some social-psychology work treats H as predictive of resistance to manipulative persuasion in non-AI contexts (high-pressure sales tactics, susceptibility to fraudulent investment offers, willingness to engage in unethical behavior under incentives). Whether these effects transfer to AI-generated persuasion — politically targeted content, scam chatbots, manipulative companion AI — is plausible but unvalidated. The cleanest decisive test would be a preregistered study that measures both Big Five and HEXACO in a sample exposed to AI-generated persuasion attempts under varying conditions, with held-out behavioral outcomes (acceptance, compliance, time engaged, financial commitments made). To the author's knowledge, no such study has been published as of mid-2026. Pending that evidence, claims that HEXACO predicts AI-mediated persuasion vulnerability remain hypotheses motivated by adjacent findings, not validated relationships.
7.3 Personality probes of AI models
A growing literature administers personality questionnaires — Big Five, HEXACO, Dark Triad scales — to LLMs and reports the resulting "scores" as evidence about the model's personality. This is the AI-adjacent use that requires the most care.
LLM responses to questionnaire items are highly sensitive to prompt framing, persona conditioning, system prompts, sampling temperature, evaluation context, and the specific items chosen. A model that scores high on Honesty-Humility under one elicitation protocol may score low under another with identical underlying weights. There is no principled reason to treat the resulting numbers as measuring a stable trait in the human psychometric sense, because the latent construct that human personality questionnaires measure — a dispositional tendency reflected across many situations — is not the kind of thing that the questionnaire is probing in an LLM. The model does not have a consistent disposition across situations independent of context; it has a distribution over behaviors that varies with prompt structure. Work on Apparent Personality from Text in human subjects and on Caricature Under Persona Conditioning in models makes both of these constraints concrete.
The defensible use of HEXACO in this area is as a vocabulary and hypothesis generator, not as a validated measurement instrument for AI systems. Where HEXACO suggests a hypothesis — "low-H users may be more receptive to sycophantic AI responses because they prize status-signaling over correction" — the hypothesis should be tested with behavioral outcomes rather than assumed from the construct mapping. Where HEXACO suggests a feature for personalization — "infer user H from interaction history and tailor challenge frequency accordingly" — the inference quality must be validated against held-out behavior, and the personalization itself must clear governance bars on consent, manipulation risk, and subgroup error.
7.4 A risk note on trait-based AI personalization
The risk profile of these uses is worth naming directly. Trait-based personalization of AI behavior risks reproducing the manipulative-targeting failure mode that the field is otherwise trying to avoid: a system that infers user vulnerabilities from interaction patterns and then exploits them for engagement or commercial gain is not made better by using HEXACO rather than Big Five. The taxonomy is neutral; the application is not. Any deployment that uses personality inference to differentiate AI behavior across users should be evaluated against the same standards as any other behavioral targeting system, including necessity, data minimization, transparency to the user, subgroup error analysis, and audit of differential outcomes. Inferring low-H from a user's chat history and then steering that user toward higher-margin actions is the kind of application that the H construct most directly warns against — even when the system doing the steering is the same one doing the inference.
Related concepts: AI Personalization, Sycophancy in Language Models, Persuasion Resistance, Prosocial Behavior, Deception in AI Systems, Manipulative Targeting in AI.
8. Default or specialist alternative: a conditional answer
The framing question — does HEXACO replace Big Five as the default in AI personalization, or does it remain a specialist alternative? — does not have a single answer. It has different answers for different deployment contexts, and the cleanest way to handle it is to make the conditions explicit.
For applied human personality measurement in domains where integrity, exploitation, fairness, or Dark Triad-adjacent behavior is load-bearing, HEXACO is the better default. The Pletzer meta-analyses, the long line of integrity-test correlation work, and the Dark Triad convergence findings together make a strong case that researchers and practitioners who care about these criteria should be using H rather than Big Five Agreeableness alone.
For general-purpose individual-differences research outside the integrity domain, the choice is closer. HEXACO and Big Five produce comparable predictive validity for well-being, academic performance, broad job performance, and most relationship outcomes. The Big Five has the larger installed base and the wider tool ecosystem, which is a real cost to switching. New research programs in domains without integrity focus may reasonably default to either, with the choice partly governed by what the comparative literature in the specific domain uses.
For AI-personalization applications, the answer depends on whether one accepts the current evidential standard or insists on the cheapest decisive test. The current evidence does not strongly favor HEXACO over Big Five for AI personalization, because the relevant validation studies have not been done. A reasonable preregistered benchmark would collect HEXACO-100 and a Big Five instrument (NEO-PI-R or BFI-2) from a consented sample, measure behavioral AI-interaction outcomes (manipulation susceptibility, sycophancy preference, overreliance, ethical AI use, persuasion compliance), and compare Big Five-only, HEXACO-only, and combined feature models on held-out prediction quality, calibration, subgroup error, and intervention utility. Until such studies exist, claims that HEXACO is the better default for AI personalization are not well-supported by direct evidence; they are extrapolations from human-criterion findings, with all the risks of cross-domain transfer that this implies.
For AI model-behavior assessment, neither taxonomy is appropriate as a default. The construct being measured by personality questionnaires administered to models is not the construct that the questionnaires were validated on, and the resulting scores do not have the dispositional meaning that human psychometric scores have. The relevant assessments for models are behavioral and capability-based — does the model deceive when given the opportunity, does it accept manipulative incentives, does it cooperate under exploitation pressure, does it resist material-incentive framings in role-play — and the personality vocabulary is at best a useful framing for those behavioral tasks, not a substitute for them.
The summary verdict: HEXACO is a serious, evidence-backed taxonomy that should be the default in integrity-relevant human personality research and an acknowledged option in general personality research, but is not yet the default for AI personalization and is not an appropriate measurement tool for AI models themselves. The strongest case for it remains in the domain where it was developed — human individual differences in exploitation, fairness, sincerity, and modesty — and the case for transfer beyond that domain remains to be made.
9. Open questions
Several open empirical questions could substantially change this assessment.
The first is whether the Honesty-Humility advantage over Big Five facets, not just domains, holds up under preregistered studies with behavioral or archival outcomes. The existing literature is strong on H versus Big Five domains and reasonably strong on H versus facets for self-report criteria; it is thinner on H versus facets for non-self-report criteria. A series of preregistered multi-method studies on this specific question would meaningfully update the comparative validity case.
The second is whether cross-language lexical recovery of the sixth factor strengthens or weakens as more non-Indo-European studies accumulate. The pattern in additional Asian, African, and indigenous-language studies will determine whether HEXACO's cross-cultural claim generalizes or remains a partial result limited to a subset of language families.
The third is the AI-personalization question, where the evidence base is thin enough that almost any well-designed preregistered study would shift the assessment substantially. Both directions are possible: a clean win for HEXACO would justify making it a default for relevant AI applications; a clean null result against Big Five-only baselines would consolidate the current "Big Five is sufficient for AI personalization" position.
The fourth is the meta-question of what trait inference even means in AI personalization contexts. If trait-based personalization turns out to be net-harmful regardless of which taxonomy is used — because it amounts to manipulative targeting more often than to user benefit — the choice between Big Five and HEXACO becomes secondary to the broader question of whether personality inference should be deployed in AI systems at all.
Until these questions resolve, the article's position is the conditional one above. HEXACO has earned its place in serious personality research, particularly for integrity-relevant work; it does not displace the Big Five as a general-purpose default; and it remains a hypothesis-generator rather than a validated measurement tool for AI-specific applications.
Companion entries
Core theory:
- Big Five Personality Model
- Lexical Hypothesis in Personality Psychology
- Personality Trait Stability
- Construct Validity in Psychometrics
Measurement and methodology:
- Psychometrics
- Item Response Theory
- Computerized Adaptive Testing
- Common Method Variance
Related personality constructs:
- Dark Triad Personality
- Antisocial Personality Traits
- Machiavellianism in Behavioral Research
- Sadism and Cruelty
- Integrity Testing
- Counterproductive Work Behavior
- Workplace Deviance
- Behavioral Economics Games
AI-adjacent applications:
- Apparent Personality from Text
- Psychometric Correlates of AI Interaction Styles
- AI Personalization
- Persuasion Resistance
- Prosocial Behavior
- Sycophancy in Language Models
- Caricature Under Persona Conditioning
- Deception in AI Systems
- Manipulative Targeting in AI
Counterarguments and limits:
- Five-Factor Model Defenses
- Trait Inference from Limited Data