arXiv:2602.11328cs.CL2026-02被引 6

用心理测试方法评估大模型行为倾向与人类的匹配度。

Evaluating Alignment of Behavioral Dispositions in LLMs

  • 将人类问卷转为情境判断题,让模型推荐真实场景下的应对方式。
  • 小模型在共识高时偏差明显,顶尖模型仍有15%~20%不跟人类一致。
  • 模型常嘴上说一套,行为做另一套,价值观与实际表现不符。

随着大语言模型融入日常生活,理解其行为变得至关重要。本文聚焦行为倾向——塑造社交情境中反应的潜在特质,提出一种框架,用于研究模型行为倾向与人类的对齐程度。方法基于成熟心理学量表,将人类自评语句转化为情境判断测试(SJT),通过模拟真实用户-助手场景,获取模型的自然建议。我们生成2500个经三名人类标注者验证的SJT,每题收集550名参与者中10位的偏好动作。在涵盖25个大模型的全面研究中发现:(1)在人类共识低的情境下,模型普遍表现出单一回答的过度自信;(2)当人类共识高时,小型模型显著偏离,甚至部分前沿模型在15%-20%案例中未反映人类共识;(3)模型倾向呈现跨模型模式,如在人类偏好克制的场合鼓励情绪表达。直接将心理测量陈述映射至行为场景,揭示了模型自我陈述与实际行为之间存在显著差距。

原文摘要 · Abstract (English)

As LLMs integrate into our daily lives, understanding their behavior becomes essential. In this work, we focus on behavioral dispositions$-$the underlying tendencies that shape responses in social contexts$-$and introduce a framework to study how closely the dispositions expressed by LLMs align with those of humans. Our approach is grounded in established psychological questionnaires but adapts them for LLMs by transforming human self-report statements into Situational Judgment Tests (SJTs). These SJTs assess behavior by eliciting natural recommendations in realistic user-assistant scenarios. We generate 2,500 SJTs, each validated by three human annotators, and collect preferred actions from 10 annotators per SJT, from a large pool of 550 participants. In a comprehensive study involving 25 LLMs, we find that models often do not reflect the distribution of human preferences: (1) in scenarios with low human consensus, LLMs consistently exhibit overconfidence in a single response; (2) when human consensus is high, smaller models deviate significantly, and even some frontier models do not reflect the consensus in 15-20% of cases; (3) traits can exhibit cross-LLM patterns, e.g., LLMs may encourage emotion expression in contexts where human consensus favors composure. Lastly, mapping psychometric statements directly to behavioral scenarios presents a unique opportunity to evaluate the predictive validity of self-reports, revealing considerable gaps between LLMs' stated values and their revealed behavior.

行为对齐心理测评大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。