Wright · Nature human behaviour 2026 · Observational validation study · n=?

Assessing personality using zero-shot generative AI scoring of brief open-ended text.

Level 3 - non-randomized controlled study

Observational validation study across two cohorts comparing AI scoring against established psychometric measures.

PubMed 41617861 · doi:10.1038/s41562-025-02389-x · record verified 2026-08-26

What was done

Across two distinct participant cohorts and data-collection modalities (spontaneous streams of thought and daily video diaries), researchers evaluated the ability of seven commercial generative large language models (including ChatGPT and Claude) to score Big-Five personality traits from brief open-ended qualitative narratives using zero-shot prompting. Trait scores derived from the models were compared against standard self-report personality measures and evaluated for predictive validity regarding daily behaviors and mental health outcomes.

What was found

The abstract reports no numeric values, effect sizes, correlation coefficients, or sample sizes. Qualitatively, zero-shot LLM trait scoring achieved convergence with self-report measures comparable to or exceeding established benchmarks such as self-other agreement, ecological momentary assessment, and bespoke machine learning models. Averaging scores across the seven LLMs provided the strongest agreement with self-report, and LLM-derived scores demonstrated predictive validity for mental health outcomes and daily behaviors.

Why it matters

This study suggests that off-the-shelf commercial LLMs can extract structured personality trait metrics directly from unstructured, qualitative narratives without task-specific fine-tuning or specialized training, potentially bridging qualitative depth and quantitative scale in psychological assessment.

Limits

The abstract provides no sample sizes, demographic characteristics, effect sizes, or confidence intervals. Relying on closed, proprietary commercial LLMs poses reproducibility challenges due to frequent unannounced model updates, and performance across diverse languages, cultures, or clinical populations remains uncharacterized.