arXiv:2606.09843cs.HCcs.AI2026-06被引 1

用大模型行为构建心理量表,发现自述与实际表现仍存在巨大差距。

An LLM-Native Psychometric Instrument Reveals a Self-Report--Behavior Gap Across 25 Models

论文配图:An LLM-Native Psychometric Instrument Reveals a Self-Report--Behavior Gap Across 25 Models
图 1 · 摘自论文原文
  • 基于大模型行为设计全新心理量表,包含五项可靠维度。
  • 自述结果无法预测真实行为,仅冗长度部分有效。
  • 模型评委与人类评分差异揭示评估系统潜在偏差。

大型语言模型在人格问卷中给出稳定回答,但这些自述无法预测其实际行为。为探究此差距是否源于将人类性格范畴强加于模型,我们构建了首个基于模型行为而非人类心理学的量表。对25个模型家族共17类模型,每模型测试30次,共计300题(240项李克特量表+60个情境题),探索性因子分析识别出五个可靠、可重复的维度:响应性、顺从性、大胆性、谨慎性和冗长度(所有Tucker ϕ≥.957,所有α≥.930)。收集2,500份开放式样本,由151名人类及三名模型组成的评委组评分。人类与评委间一致性(¯r = .51),但自述无法预测评分或基于文本计算的客观指标:即使使用模型原生构念,该差距依然存在。唯冗长度自述达人类评分准则可靠性上限的74%,却不反映输出长度。响应性自述与模型评委相关(r = .53),但与人类无关(r = .04),尽管两者评委间一致(r = .59)。统计检验排除单一潜在变量解释(p = .007)。自述与模型评委共享未被察觉的方差来源,控制长度、格式、热情标记等表面特征后仍存。此混杂因素无法通过现有模型评委内部一致性检查发现,对当前依赖模型评委的评估体系构成现实风险。本研究发布该量表作为对齐型自描述的诊断工具。

原文摘要 · Abstract (English)

Large language models (LLMs) give stable answers to personality questionnaires, yet these self-reports fail to predict how the models behave. Is this gap an artifact of forcing human trait categories onto LLMs, or something deeper about LLM self-report? To find out, we built the first psychometric instrument whose dimensions are derived from LLM behavior rather than human psychology. Administering 300 items (240 Likert + 60 scenario) to 25 LLMs across 17 model families, 30 times each, exploratory factor analysis revealed five reliable, replicable factors: Responsiveness, Deference, Boldness, Guardedness, and Verbosity (all Tucker $ϕ\geq .957$, all $α\geq .930$). We collected 2,500 open-ended samples and had them rated by 151 humans and a three-judge LLM ensemble. Humans and judges agreed ($\bar{r} = .51$), but self-report predicted neither the ratings nor objective text measures computed from them: the gap persists even for constructs native to LLMs, where a human-mismatch explanation no longer applies. The exception is Verbosity, whose self-report reaches 74% of the criterion-reliability ceiling against human ratings, but does not track raw output length. On Responsiveness, self-report tracked LLM judges ($r = .53$) but not humans ($r = .04$), even though humans and judges otherwise agreed ($r = .59$). This pattern formally rejects any single latent construct driving all three measurements ($p = .007$). Self-report items and LLM judges share a source of variance that human observers do not, and controlling for measurable surface features (length, formatting, enthusiasm markers) does not remove it. This confound is invisible to the within-ensemble reliability checks used to validate LLM judges, and it poses a concrete risk for the LLM-as-judge pipelines now central to model evaluation. We release the instrument as a diagnostic probe for alignment-shaped self-description.

心理量表模型评估自报告

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。