对齐让模型更符合规范,反而降低对真实人类行为的预测能力。
Alignment Makes Language Models Normative, Not Descriptive
- 对比基础模型与对齐模型在真实人类决策中的表现
- 多轮策略游戏中,基础模型预测准确率高出对齐模型近10倍
- 对齐模型仅在单次教科书游戏和非策略选择中表现更好
后训练对齐通过优化语言模型以匹配人类偏好信号,但该目标并不等同于建模实际观察到的人类行为。我们比较了120组基线对齐模型对超过10,000个真实人类决策的表现,涵盖多轮策略博弈——讨价还价、说服、谈判和重复矩阵博弈。在这些场景中,基线模型在预测人类选择上的表现比对齐模型高出近10:1,且结果在不同模型族、提示形式和游戏配置下均稳健。然而,在人类行为更趋近规范性预测的情境中,对齐模型则占优:在所有12种单次教科书博弈及非策略性彩票选择中表现优异;甚至在多轮博弈的第一轮(互动历史尚未形成前)也表现更好。这一边界条件模式表明,对齐引入了规范性偏差:当人类行为较易被规范解解释时,对齐模型预测更准;但在多轮策略场景中,人类行为受互惠、报复和历史依赖适应等描述性动态影响,此时对齐模型反而预测能力下降。研究揭示了优化模型用于人类使用与将其作为人类行为代理之间的根本权衡。
原文摘要 · Abstract (English)
Post-training alignment optimizes language models to match human preference signals, but this objective is not equivalent to modeling observed human behavior. We compare 120 base-aligned model pairs on more than 10,000 real human decisions in multi-round strategic games - bargaining, persuasion, negotiation, and repeated matrix games. In these settings, base models outperform their aligned counterparts in predicting human choices by nearly 10:1, robustly across model families, prompt formulations, and game configurations. This pattern reverses, however, in settings where human behavior is more likely to follow normative predictions: aligned models dominate on one-shot textbook games across all 12 types tested and on non-strategic lottery choices - and even within the multi-round games themselves, at round one, before interaction history develops. This boundary-condition pattern suggests that alignment induces a normative bias: it improves prediction when human behavior is relatively well captured by normative solutions, but hurts prediction in multi-round strategic settings, where behavior is shaped by descriptive dynamics such as reciprocity, retaliation, and history-dependent adaptation. These results reveal a fundamental trade-off between optimizing models for human use and using them as proxies for human behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。