Peters · PNAS nexus 2024 · cross-sectional computational validation study · n=?

Large language models can infer psychological dispositions of social media users.

Level 4 - case-series / case-control

Cross-sectional computational validation study comparing LLM predictions against self-reported ground truth

PubMed 38948324 · doi:10.1093/pnasnexus/pgae231 · record verified 2026-08-26

What was done

Researchers evaluated the zero-shot capability of GPT-3.5 and GPT-4 to infer Big Five personality traits from Facebook status updates. LLM-derived trait predictions were compared against users' self-reported scores, and accuracy was analyzed across age and gender subgroups.

What was found

LLM-inferred personality scores correlated with self-reported traits with an average correlation of r = 0.29 (range: 0.22 to 0.33). Accuracy varied across demographics, with predictions being more accurate for women and younger individuals across several traits.

Why it matters

This shows that general-purpose LLMs can infer psychological traits from unstructured online text without task-specific training, introducing opportunities for low-cost psychometrics as well as significant privacy and profiling risks.

Limits

The abstract does not state the sample size (n), participant demographics, or the specific number of status updates analyzed. The observed correlations (r = 0.22 to 0.33) leave substantial variance unexplained, and performance differed across demographic groups.