arXiv:2606.12730cs.AIcs.CL2026-06中稿 · as an Oral被引 1

用行为意图理论比性格测试更能预测大模型行为。

Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior

论文配图:Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior
图 1 · 摘自论文原文
  • 改用行为意图理论替代大五人格评估模型行为
  • 同一对话中意图预测达人类水平,跨对话仅在隐性偏见上有效
  • 角色提示提升自述一致性但未对齐实际行为

从低成本心理测量探针预测大模型行为倾向对安全部署至关重要,前提是自述(SR)能可靠预测行为。近期研究发现大模型存在自述与行为的显著分离,但其使用的大五人格等宽泛特质在人类中也难以预测具体行为。此外,对话会话孤立且上下文匹配弱,无法判断模型是否真正缺乏连贯性。本文对比大五人格与计划行为理论(TPB),后者针对特定行为测量意图,预测能力远超宽泛特质。在四个行为任务和11个前沿大模型上实验,同时改变会话上下文与身份诱导条件。结果表明:1)同一对话中,计划行为理论达到人类水平的一致性,大五人格则不行;2)跨对话时,仅当行为由训练数据塑造的隐性偏见锚定才保持一致,若行为受上下文强烈引导(如迎合)则迅速崩溃;3)角色提示使自述在不同对话中更一致,但行为仍不对其齐。结论指出,大五等粗粒度框架不适合作为部署行为测试工具,需采用任务和行为特异性强的测量手段,并在多任务与上下文中验证。

原文摘要 · Abstract (English)

Anticipating LLM behavioral tendencies from low-cost psychometric probes is critical for safe deployment, but only if self-reports (SR) reliably predict behavior. Recent work documented substantial SR-behavior dissociation in LLMs, but relied on broad personality traits (Big 5) that predict specific behaviors weakly, even in humans. Furthermore, the isolation of conversational sessions combined with weak context matching left open whether LLMs truly lack coherence or whether the conditions needed to detect such coherence were not met. We contrast Big 5 with the Theory of Planned Behavior (TPB), which measures intention targeted to a specific behavior and predicts human behavior substantially better than broad traits. We run experiments across four behavioral tasks and 11 frontier LLMs, while also varying session context and identity induction. We find that SR-behavior coherence exists but is selective. 1) Within a shared conversation, the Theory of Planned Behavior reaches human-level coherence; Big 5 does not. 2) Across separate conversations, coherence survives only for behaviors anchored outside the immediate prompt, such as implicit bias shaped by training, and collapses when behavior is strongly primed by context, as with sycophancy. 3) Persona prompting makes self-reports more consistent across conversations, but does not bring behavior into alignment. These findings suggest that coarse personality frameworks, such as Big 5 may not be the best tools for testing deployment behavior. More task- and behavior-specific instruments are needed, and even these must be evaluated across tasks and contexts.

大模型评测行为预测心理测量计划行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。