Ashita Orbis

Apparent Personality from Text: LLM-Based Personality Inference

This article covers the line of work that maps written text to Big Five-like personality scores, from Mairesse et al.'s handcrafted linguistic features through open-vocabulary social-media studies to recent LLM-prompted assessment. The empirical pattern across two decades is consistent: language carries real personality-related signal, and recent LLM-based methods can recover that signal under specific elicitation conditions — but the headline correlations between LLM-inferred scores and self-report cluster in the .3-.45 range, the prompt is part of the measurement instrument, and the construct being recovered is closer to apparent personality from a situated text sample than to direct access to latent traits.

Coverage note: verified through May 2026.

1. The measurement question, before the methods

Most of the confusion in this area is a target problem, not a method problem. "Personality" in the literature actually refers to at least four distinct criteria that get traded for one another in the same paragraph:

  1. Actual personality — a hypothesized latent disposition that influences behavior across situations.
  2. Self-report Big Five — scores from a standardized questionnaire (BFI-2, IPIP, NEO-FFI), which are the operational benchmark in nearly all text-personality work.
  3. Informant report — Big Five ratings produced by someone who knows the target, typically with modest agreement to self-report (single-informant agreement is usually in the r=.40-.55 range for the Big Five).
  4. Apparent or expressed personality — how a target appears from a defined behavioral slice, modeled in the apparent-personality and first-impressions literature (Ponce-López et al., 2016; Vinciarelli & Mohammadi, 2014).

Text inference can address any of these targets, but it does not get to choose silently. A study that scores Facebook statuses against self-report is producing convergent evidence between two methods on partially overlapping construct space; it is not measuring latent personality, and it is not the same as the apparent-personality task in which independent raters produce trait impressions from expressive behavior.

The risk that runs through this field is construct laundering: a chain that begins with "model output correlates with self-report at r≈.4," passes through "model infers personality," and arrives at "model has measured the user's personality," with the criterion silently swapped at each step. The article that follows organizes the evidence so the criterion stays attached to the claim. Related: Construct Validity in Psychometrics, Apparent Personality, Big Five Personality Model.

2. The lineage from handcrafted features to LLM prompts

2.1 Mairesse, Walker, Mehl, Moore (2007): handcrafted linguistic cues

The first systematic attempt to recognize personality from text used dictionary-based features (LIWC, MRC psycholinguistic database) and prosodic cues, with regression and classification models trained on conversation transcripts (the EAR corpus) and written essays (the Pennebaker stream-of-consciousness corpus). Mairesse et al. (2007) reported that observer-rated personality was generally easier to model than self-reported personality, that extraversion was the most recoverable trait, and that even simple feature families produced above-baseline performance. Two features of that result still matter:

  • It already separated self-report and observer-rated criteria and showed that observer judgments were the more tractable target. That is a near-direct anticipation of the apparent-personality frame.
  • It established the empirical ceiling pattern that has not really moved: trait-by-trait correlations in the .15-.40 range against self-report, with extraversion most predictable and openness somewhat less so.

2.2 Schwartz et al. (2013): open-vocabulary differential language analysis

Schwartz and colleagues replaced the closed dictionary with an open-vocabulary approach: they extracted unigrams, n-grams, and topics (LDA-derived) from 700M words of Facebook status updates produced by 75,000 myPersonality volunteers and looked at differential frequency by Big Five trait, age, and gender (Schwartz et al., 2013). The methodological contribution was twofold. First, the word clouds were psychologically informative in a way LIWC categories could not be (the high-extraversion cluster included party-related n-grams, the high-neuroticism cluster included disclosure and complaint terms, the high-conscientiousness cluster included scheduling and effort language). Second, the demographic signals (age, gender) were at least as strong as the trait signals, which is part of why later corpus-transfer failures are unsurprising.

The Schwartz paper is also where the field's reliance on Facebook status updates as the canonical corpus began. That choice has costs that are still being paid: status updates are short, public, performative, and skewed toward certain age/education ranges. The findings are best read as "language used in this corpus differentiates people who score differently on a self-report Big Five questionnaire," not "language differentiates people."

2.3 Park et al. (2015): the strongest psychometric case for language-based assessment

Park et al. (2015) trained a language-based assessment on 66,732 Facebook users and validated on a held-out 4,824, comparing the language-derived scores against self-report Big Five, informant report, external life-outcome criteria, and six-month test-retest stability of the language-derived scores themselves. This is the most psychometrically serious paper in the pre-LLM line. Two numbers matter.

  • The average convergence with self-report was about r=.38 across the five domains. That is the moderate-correlation ceiling that recurs across the literature.
  • The language-derived scores converged with informant ratings about as well as informants converged with each other, and they predicted some life outcomes (life satisfaction, depression) at levels comparable to self-report.

Park et al. is the result most commonly invoked when defending text-based personality assessment as a serious measurement modality. Read carefully, what it supports is that language-based scores behave like another method of measuring Big Five-like constructs, with method variance and bias characteristic of its corpus, rather than a transparent window onto traits.

2.4 The LLM turn: 2023-2024

Between 2023 and 2024 the methodology shifted from supervised models trained on labeled language to prompted LLMs producing scores directly. Three contributions define the current state.

Rao, Leung & Miao (2023) is the earliest widely cited "Can ChatGPT assess human personalities" framework. The paper proposes a general MBTI-based evaluation protocol in which ChatGPT scores written and dialogic input against trait dimensions and reports that the model's assessments are more consistent and fairer than human raters in some conditions, with lower robustness to perturbations. The MBTI framing is a substantive limitation — MBTI is a weaker construct than the Big Five and not interchangeable with it — but the paper's methodological contribution (an evaluation harness for LLM-as-personality-rater) is genuine. (The original research prompt for this article also referenced a "Yang et al. 2023" ChatGPT personality assessment paper; the citation most clearly identifiable in the EMNLP-2023 LLM-personality literature is Rao, Leung & Miao. If a separate Yang et al. paper is intended, it should be cited specifically rather than under a generic author tag.)

Peters & Matz (2024) is the larger and more rigorous LLM-personality study in the social-media line. The authors had GPT-3.5 and GPT-4 score Big Five traits zero-shot from the Facebook statuses of myPersonality users and compared against self-report. The mean correlation across the five traits was r=.29 (range .22-.33), comparable to supervised lexical models trained explicitly for the task. The paper is also one of the few in this corner of the literature that reports demographic-subgroup variation: predictions were systematically more accurate for women and younger users on several traits, which is exactly the kind of bias signal that earlier text-personality work tended not to report.

Peters, Cerf & Matz (2024) moved the question from static social-media text to interactive chat. They had GPT-4 chatbots converse with users under three elicitation conditions, then score the participants' Big Five from the transcript:

  • An assessment-optimized prompt instructed the model to elicit personality-relevant information conversationally. Mean correlation with self-report: r=.443, range .245-.640.
  • A naturalistic-interaction prompt prioritized comfortable conversation over assessment. Mean correlation: r=.218, range .066-.373.
  • A default helpful-assistant condition produced the weakest correspondence with self-report.

This is the most important LLM-personality result for systems that ingest dialogue. It shows that the headline number this field uses to argue for or against LLM personality inference is a function of the prompt, the elicitation context, and the implicit goal placed on the model — not a property of "GPT-4" or "language models." The prompt is part of the instrument.

2.5 Smartphone-sensing as a comparison class

Stachl et al. (2020) is the right non-text reference point. The authors used 30 days of passively collected smartphone behavior (communication patterns, app usage, mobility, music, day-night activity) on 624 participants to predict Big Five domains and facets. Median cross-validated correlation for domains was r=.37; for facets, around r=.40. Communication and social behaviors were the most predictive feature family.

The Stachl result is useful for two reasons. First, it independently lands on the same moderate ceiling (~r=.4) from a completely different data modality, which is evidence that the ceiling is a property of the criterion (self-report Big Five) and the prediction problem (recovering trait-level signal from situated behavior), not of any one method. Second, it makes the privacy and infrastructure costs of behavioral personality inference legible: 30 days of passive sensing is a heavy infrastructure ask for a correlation that is similar to what a few thousand words of conversation can produce under good prompting.

3. Effect sizes and what they actually license

The following table tracks where each major result sits in claim space.

Study Text/data source Criterion Method Reported strength Claim licensed
Mairesse et al. 2007 Conversation transcripts, essays Self-report and observer-rated Handcrafted features + regression Trait-level r in .15-.40, observer > self Linguistic cues carry trait signal; observer-rated is the easier target
Schwartz et al. 2013 700M words of Facebook status myPersonality self-report Big Five Open-vocabulary differential Population-level trait/word correlations Language usage differentiates self-reported trait groups in this corpus
Park et al. 2015 Facebook status; n≈66k train, ~5k validate Self-report, informant, outcomes Supervised regression on language Mean r≈.38 to self-report; informant parity Language-based scores behave as a method of Big Five measurement with own bias
Peters & Matz 2024 Facebook statuses, myPersonality Self-report Big Five GPT-3.5/GPT-4 zero-shot prompting Mean r=.29 (range .22-.33); subgroup bias Zero-shot LLMs match supervised lexical baselines on this corpus, with bias
Peters, Cerf & Matz 2024 GPT-4 chatbot dialogue, n=722 Self-report Big Five Assessment-optimized vs. naturalistic prompts Mean r=.443 / .218 / weaker for default helper Elicitation prompt is part of the instrument; effect size is prompt-conditional
Rao, Leung & Miao 2023 Mixed text and dialogue MBTI-style trait categories ChatGPT zero-shot scoring Consistency over individual humans; low robustness LLMs can produce trait ratings; MBTI framing limits transfer to Big Five
Stachl et al. 2020 (non-text) 30-day smartphone sensing Self-report Big Five (BFSI) Cross-validated ML Median domain r=.37, facet r=.40 Independent ceiling from a different modality; communication signals strongest

A few things follow from reading this table sideways.

First, the moderate-correlation ceiling around r=.3-.4 against self-report holds across two decades, four method families (handcrafted lexical features, open-vocabulary supervised models, zero-shot LLM prompting, behavioral sensing) and two large independent corpora (myPersonality Facebook, smartphone-sensed behavior). That convergence is meaningful: it suggests the ceiling is driven by criterion noise, construct mismatch between method and target, and the limits of recovering individual trait scores from situated behavior — not by any specific algorithmic weakness.

Second, the r=.443 headline from Peters, Cerf & Matz is at the upper end of that pattern, achieved under elicitation conditions designed to extract trait-relevant information. The same architecture with a default helpful-assistant prompt falls to roughly r=.117. A field that reports the best of those numbers without the prompt context is reporting a method, not an instrument property.

Third, the population-level versus individual-level distinction matters more than it usually gets credit for. A correlation of r=.4 is a serious effect at the level of group comparisons and aggregate prediction. It is a much weaker basis for individual trait classification, where the squared correlation (about .16) gives a rough sense of how much individual-level variance the inference actually explains. For most personalization or screening use cases, "explains 16% of self-report variance" is the right number to think about, not the bare correlation.

4. The prompt is the instrument

The strongest practical takeaway from the LLM line is that prompt design, elicitation context, and scoring rubric belong inside the measurement instrument, not in a methods footnote. Three sub-patterns have emerged.

Assessment-optimized prompts. Peters, Cerf & Matz's strongest condition explicitly told the model to elicit personality-relevant information conversationally. This recovers far more signal than a default assistant prompt because it changes what the model attends to in the user's text and what kind of follow-ups it asks. Conceptually, this is identical to the difference between an unstructured conversation and a clinical intake — different elicitation methods produce different evidence.

Few-shot examples. Providing the model with example transcripts annotated with trait scores tends to improve calibration but also imports the rater conventions of whoever produced the examples. The model is fitting the scoring rubric of the demonstration set, not converging on an external truth about traits.

Chain-of-thought-style scoring. Asking the model to lay out its reasoning trait-by-trait before producing a numeric score sometimes improves correlations with self-report. It also introduces a new failure mode: the model may produce plausible-sounding trait narratives that exceed the evidence in the text, then anchor the numeric score on the narrative rather than on the linguistic evidence. The reasoning trace looks more interpretable than a single number, but its faithfulness to the actual basis of the score is not guaranteed. Related: Chain-of-Thought Faithfulness.

The right way to treat these variants is as distinct measurement procedures that share a backbone model. Comparing "LLM personality inference" across labs without naming the prompt class, scoring rubric, model version, temperature, and elicitation context is comparing different instruments. Preregistered prompts, frozen scoring rubrics, and held-out validation are minimum hygiene if the field wants to interpret prompt differences as anything other than researcher degrees of freedom.

5. Construct validity: apparent versus actual, and what self-report is good for

The cleanest conceptual move is to relabel the modality. "LLM-based personality inference" sounds like trait extraction; "apparent personality from text" is closer to what the evidence supports. The apparent-personality literature (Ponce-López et al., 2016; Vinciarelli & Mohammadi, 2014) was developed for exactly the methodological situation that LLM inference now occupies: a model or judge produces trait impressions from a defined behavioral slice, those impressions are reliable across raters, and they correlate moderately with self-report or external criteria. LLMs are not new in this picture; they are an unusually scalable, culturally competent artificial rater.

This framing is not a demotion. It clarifies what the modality is good for and what it is not.

What apparent-personality measurement is good for:

  • Capturing the expressed persona in a written corpus (an author's blog voice, a chat user's communicative style, a candidate's writing sample). That is exactly what readers and collaborators perceive, and for many downstream questions (writing tone, communication fit, content moderation, interview screening) it is the directly decision-relevant variable.
  • Population-level inferences over text-rich data (longitudinal pulse-of-platform analysis, content-region comparisons, hypothesis generation for follow-up self-report studies).
  • Triangulation alongside other measurement methods.

What apparent-personality measurement does not deliver:

  • Direct access to latent traits independent of corpus and context.
  • Individual-level classification with the confidence that a clinical or hiring decision usually requires.
  • Robustness to topic, genre, demographic, or platform shifts without explicit revalidation.

Self-report is the criterion the field uses, but it is not ground truth in the strong sense. Self-report Big Five contains response-style variance, self-presentation effects, item-interpretation differences, limited introspective access, and stable demographic patterns. The correct claim about a study reporting r=.4 with self-report is that two methods (the text-derived score and the questionnaire) converge moderately on a partially shared construct. Treating self-report as a stand-in for "actual personality" inflates the conclusion in one direction; treating the LLM as a direct trait reader inflates it in the other.

Where the field is short on evidence is in convergence with informant ratings, behavioral predictions, and future outcomes from text alone, especially under preregistered prompts. Park et al. 2015 did the most serious work on this for the supervised-lexical era; the analogous study has not yet been done at scale in the LLM-prompted era.

6. Confounds the field has not solved

6.1 Corpus selection: published writing is not spoken word

The corpus matters as much as the model. Facebook status updates, chat transcripts with an assessment-optimized agent, polished blog posts, technical writing, therapy-style reflections, customer-service messages, and spoken transcripts are not interchangeable language samples. They differ in:

  • Audience design: the writer's anticipated reader shapes what gets disclosed.
  • Editing time: edited writing tends to look more conscientious and emotionally stable than spontaneous speech, regardless of the writer's traits.
  • Genre norms: technical prose suppresses affective and self-referential language; therapeutic prose amplifies them.
  • Topic constraints: someone writing about grief or illness will look more neurotic than the same person writing about a hobby.

The Peters & Matz r=.29 number is from Facebook statuses. The r=.443 number is from a dialogue specifically designed to elicit personality-relevant content. Generalizing either to "a user's text" without naming the genre is the most common overclaim in deployment-oriented summaries of this literature.

6.2 Demographic and topic leakage

Schwartz et al. already showed that age and gender signals in language can be as strong as trait signals. Peters & Matz found subgroup variation in inference accuracy. Language encodes class, education, native-language status, profession, region, platform norms, and culture in ways that are easy to mistake for trait variance. A model that infers conscientiousness partly from polished grammar, openness partly from elite cultural references, or extraversion partly from topic frequency may be tracking the social conditions of writing rather than the traits.

This is not solved by "controlling for demographics" in a regression — by the time demographics correlate with linguistic features that correlate with self-report scores that themselves correlate with demographics, the partialling is doing work the researcher cannot easily audit. The right move is to report subgroup residuals (does the model under- or over-predict each trait for each major subgroup?), audit them, and refuse to deploy at the individual level until the residuals are tolerable for the decision context.

6.3 State versus trait

Traits are by definition cross-situational and stable. Text is situational. A user writing in the middle of a personal crisis will produce text that looks high-neuroticism and low-agreeableness for reasons that have nothing to do with their stable disposition. A user writing in a celebratory mood will look extraverted. The Big Five framework was developed against questionnaires designed to average over situations; text-based inference operates on a single situation and silently treats it as a slice of the trait.

In practice this matters most when text volume is small, the genre is emotionally charged, or recent state is unusual (illness, bereavement, deadline pressure, conflict). Without a state filter, an inference layer can produce a trait label that locks in a state.

6.4 Benchmark contamination

The myPersonality dataset, the Pennebaker essay corpus, and the EAR corpus have been described in the literature for over a decade. LLMs trained on broad web text plausibly have at least partial exposure to descriptions, summaries, and replication studies of these benchmarks. A 2024 zero-shot result on myPersonality is therefore not strictly zero-shot in the way a clean held-out evaluation would be. Validation against private, freshly collected corpora is the only honest way to ask whether an LLM-prompted scorer has learned to score language or to recognize famous datasets. Related: Benchmark Contamination.

6.5 Researcher degrees of freedom in prompt engineering

If the prompt is part of the instrument, prompt search is a form of researcher freedom. A paper that tries dozens of prompt variants and reports the best correlation, without preregistration, has inflated the effective effect size in a way that is difficult to detect from outside. The Peters/Cerf/Matz result is more defensible than most precisely because the three elicitation conditions are reported separately, which exposes the prompt-sensitivity itself as a finding. A field-wide norm of preregistered prompts, frozen rubrics, and held-out validation is the minimum required to interpret cross-paper effect-size comparisons.

7. Privacy, consent, and feedback loops

Apparent-personality inference from text has properties that other psychometric methods do not. It is passive (it does not require the subject to take an instrument), it is asymmetric (a system can run inference at scale without the subject knowing), it can be applied retroactively (any past text is fair game), and it can be applied to non-consenting third parties (whoever is on the other side of a conversation, or whose writing appears in a corpus). These properties make this modality unusually well-suited to inference uses that the subject would not authorize and would not be able to detect.

Two failure modes deserve direct mention. Covert profiling uses apparent-personality inference as a third-party tool, scoring users without their participation in any explicit assessment. Feedback-loop adaptation uses an inferred trait label to alter the system's behavior toward the user, then collects the user's responses to the altered behavior as further evidence of the trait, gradually drifting away from any external check. Both failure modes are easier to build than to detect from outside. The minimum hygiene for any deployment is: explicit user-facing disclosure of inference, uncertainty display, non-sticky labels (the score changes when the evidence changes), and the user's ability to correct or suppress the inference.

8. Psyche's use of the LLM-inference layer

Psyche uses an LLM-inference layer modeled on Peters & Matz's prompting pattern as one of several inputs to a Big Five triangulation, with the LLM layer weighted at 0.25 in the combined estimate. That weight is the right place to describe what the system actually claims.

The 0.25 weight is governance, not validation. It encodes a deliberate choice that the LLM-inferred score should not be allowed to dominate other evidence — self-report instruments, longitudinal behavioral signals, and direct user correction — even when the LLM layer is internally confident. The weight functions as an uncertainty gate: the LLM contributes signal where it has signal, but its disagreement with stronger-criterion sources is resolved against it, not in favor of it.

For that weighting to mean what it is supposed to mean, several conditions must hold in the implementation.

Text-quality gates. The LLM layer should require minimum text volume and warn or downweight when the available corpus is short, single-genre, or recent. Inferring traits from 300 words of a single chat session is a different operation from inferring traits from 50,000 words across a year of varied writing, and the system should treat them as such.

Prompt versioning. Because the prompt is part of the instrument, every score should carry the prompt version, model version, and elicitation context that produced it. Switching the prompt class changes the instrument; re-scoring against a new prompt without flagging the change is silently shifting the construct.

Conflict rules. When the LLM layer disagrees materially with self-report or longitudinal behavior, the system should preserve that disagreement rather than averaging it away. The triangulation should be auditable: a profile should be able to display "LLM-derived persona suggests X, self-report suggests Y, behavior suggests Z" rather than collapsing to a single number.

Subgroup auditing. Periodic checks of the LLM layer's residuals against self-report for demographic and genre subgroups, with explicit thresholds for when to retrain the prompt or reweight the layer.

Non-sticky labels. A change in evidence should produce a change in the inferred profile, with the time-since-update visible to the user. Trait labels that persist forever after one inference run reproduce the worst failure mode of zodiacal personality systems.

The honest framing is: Psyche includes an apparent-personality-from-text layer at 0.25 weight as a weak triangulation input, calibrated against self-report under a specific prompting pattern, with explicit governance against overreach. It is not a personality test; it is a method-of-many in a multi-method profile, and its weight expresses how much the rest of the system should listen to it when other evidence is sparse.

9. The dominant-modality question

Whether LLM-based text inference becomes the dominant personality-measurement modality, or remains complementary to self-report, depends on whether the field can close gaps that the current evidence does not address.

Conditions under which dominance is plausible:

  • Text is abundant and longitudinal; self-report is unavailable or stale.
  • The decision context is low-stakes, aggregate, or reversible (research hypothesis generation, content personalization, low-friction product tuning).
  • The relevant target is expressed persona — what readers perceive — rather than latent traits. Here LLM inference is arguably already the appropriate primary instrument, because that is what it actually measures.
  • Decision processes already tolerate moderate single-criterion accuracy because they combine evidence sources downstream.

Conditions under which it should remain complementary:

  • Decisions are high-stakes and individual (hiring, clinical, legal, insurance). The 16%-of-variance figure does not support individual classification at the confidence those contexts require, and the construct concerns make even higher correlations insufficient.
  • Text is sparse, performative, or unrepresentative of the construct of interest.
  • Subgroup residuals are unaudited.
  • The system would feed back into the user's behavior in a way that confirms its own inferences.
  • Genuine latent-trait questions (longitudinal stability across decades, behavior across novel situations, clinical assessment) where established self-report and informant methods have far more validity evidence.

The cheapest decisive experiment for the field is also clear: a preregistered, multi-corpus study with a demographically balanced sample, self-report + informant + behavioral criteria, several text genres per participant (spontaneous chat, assessment-optimized chat, edited writing, transcribed spoken language), and scoring under naive prompts, assessment-optimized prompts, few-shot prompts, chain-of-thought prompts, and a strong supervised lexical baseline. The decisive outputs are cross-validated correlations by trait and corpus, incremental validity over self-report and informant report, subgroup residuals, prompt-variant robustness, and individual-level calibration error. Until that study exists, "complementary" is the empirically defensible position.

10. Bottom line

LLMs can infer Big Five-like personality scores from text with moderate convergent validity against self-report, comparable to supervised lexical models trained for the task and to behavioral inference from passive smartphone sensing. The empirical ceiling sits in the .3-.45 range across two decades and four method families, with prompt design and elicitation context capable of moving the headline by twenty correlation points or more on the same population. The construct most defensibly measured is apparent personality from a specific text sample, not latent personality; the apparent-personality methodology literature is the conceptual home, not pure trait psychometrics. For systems like Psyche that triangulate Big Five signals, the LLM-inference layer is defensible as a low-weight, uncertainty-gated input with explicit prompt versioning, subgroup auditing, and conflict rules. It is not defensible as a standalone personality test, and the field has not yet produced the cross-corpus, preregistered, demographically audited evidence that would make a "dominant modality" claim empirically supported.

The two failure modes the article most wants to prevent are linguistic. The first is the slide from "the model's score correlates with self-report" to "the model has measured personality." The second is the slide from "text inference at 0.25 weight is conservative" to "0.25 weight is automatically trustworthy." Both are failures of construct discipline, not of the underlying methods, which are useful within the limits the evidence supports.

Companion entries

Core theory:

Measurement and method:

  • Open-Vocabulary Differential Language Analysis
  • Supervised Lexical Models for Personality
  • Smartphone-Sensing Personality Validation
  • Informant-Report Personality Assessment
  • Personality Computing

LLM-specific:

Practice:

  • Psyche
  • Personality Triangulation Architecture
  • Demographic Subgroup Auditing
  • Preregistration of Prompts
  • Apparent-Persona Use Cases

Counterarguments and risks:

  • Construct Laundering
  • Covert Profiling
  • Feedback-Loop Adaptation in Personalization
  • Method Variance in Personality Measurement
  • Self-Report as Criterion, Not Ground Truth

Primary sources referenced

  • Mairesse, Walker, Mehl, Moore (2007), Using Linguistic Cues for the Automatic Recognition of Personality in Conversation and Text, Journal of Artificial Intelligence Research. (JAIR)
  • Schwartz et al. (2013), Personality, Gender, and Age in the Language of Social Media: The Open-Vocabulary Approach, PLOS ONE. (PLOS ONE)
  • Park, Schwartz, Eichstaedt, Kern, Kosinski, Stillwell, Ungar, Seligman (2015), Automatic Personality Assessment Through Social Media Language, Journal of Personality and Social Psychology. (APA PsycNet)
  • Rao, Leung, Miao (2023), Can ChatGPT Assess Human Personalities? A General Evaluation Framework, Findings of EMNLP 2023. (ACL Anthology)
  • Peters, Matz (2024), Large Language Models Can Infer Psychological Dispositions of Social Media Users, PNAS Nexus 3(6) pgae231. (PNAS Nexus)
  • Peters, Cerf, Matz (2024), Large Language Models Can Infer Personality from Free-Form User Interactions, arXiv:2405.13052. (arXiv)
  • Stachl et al. (2020), Predicting Personality from Patterns of Behavior Collected with Smartphones, PNAS 117(30) 17680-17687. (PNAS)
  • Ponce-López, Chen, Oliu, et al. (2016), ChaLearn LAP 2016: First Round Challenge on First Impressions — Dataset and Results, ECCV 2016 Workshops. (Springer)
  • Vinciarelli, Mohammadi (2014), A Survey of Personality Computing, IEEE Transactions on Affective Computing. (IEEE Xplore)

AI-researched reference article. Something wrong here? Tell us.