arXiv:2602.15848cs.CLcs.AI2026-02

用对话代替问卷,大模型可有效评估人格特质

Can LLMs Assess Personality? Validating Conversational AI for Trait Profiling

  • 通过引导式对话获取性格数据,与传统问卷对比验证
  • 尽责性、开放性和神经质得分一致,外向性和宜人性有差异
  • 用户认为大模型评估结果与传统方式一样准确

本研究验证了大型语言模型(LLMs)作为问卷法人格评估的动态替代方案。在一项被试内实验(N=33)中,我们比较了基于引导式LLM对话得出的五大性格维度评分与金标准IPIP-50问卷的结果,并测量了用户对评估准确性的感知。结果显示中等程度的一致性(r=0.38–0.58),尽责性、开放性和神经质得分在两种方法间无统计学差异;宜人性和外向性则存在显著差异,表明需进行特质特异性校准。值得注意的是,参与者对LLM生成的性格画像的准确性评价与传统问卷相当。这些发现表明,对话式AI为传统心理测量提供了一种有前景的新路径。

原文摘要 · Abstract (English)

This study validates Large Language Models (LLMs) as a dynamic alternative to questionnaire-based personality assessment. Using a within-subjects experiment (N=33), we compared Big Five personality scores derived from guided LLM conversations against the gold-standard IPIP-50 questionnaire, while also measuring user-perceived accuracy. Results indicate moderate convergent validity (r=0.38-0.58), with Conscientiousness, Openness, and Neuroticism scores statistically equivalent between methods. Agreeableness and Extraversion showed significant differences, suggesting trait-specific calibration is needed. Notably, participants rated LLM-generated profiles as equally accurate as traditional questionnaire results. These findings suggest conversational AI offers a promising new approach to traditional psychometrics.

人格评估大模型对话系统心理测量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。