arXiv:2606.22974cs.AI2026-06

LLM偏好不等于行为动力,真实任务中高偏好选项无法提升输出质量

When Preferences Fail to Become Incentives: A Utility-Behavior Gap in Large Language Models

论文配图:When Preferences Fail to Become Incentives: A Utility-Behavior Gap in Large Language Models
图 1 · 摘自论文原文
  • 通过选择实验验证模型存在一致偏好,但这些偏好未必驱动实际行为
  • 在写作、翻译等真实任务中,高偏好激励无法提升模型输出质量
  • 发现模型偏好与行为之间存在显著脱节,警示安全评估需谨慎推断

大型语言模型(LLMs)在选择范式中表现出一致的偏好结构,甚至包含训练者未意图的偏见(如某些国籍更受青睐)。然而,此类偏好是否真正驱动模型行为尚不清楚。本文设计了真实写作任务(论文、项目摘要、事故复盘、翻译),由独立的盲评LLM团队评估质量。实验表明,直接鼓励可有效调节输出质量。但当向模型提供其在选择中明确偏好的高价值结果作为激励时,其输出质量并未优于无激励或低偏好激励的情况。在所有任务和模型中均未观察到偏好激励的促进作用。结论:选择范式中显现的偏好不应被视为模型的实际动机,偏好与行为间存在显著效用-行为鸿沟。

原文摘要 · Abstract (English)

Recent work on preference elicitation in large language models (LLMs) has demonstrated that, when given a series of choices between two outcomes, LLMs reveal a coherent, model-specific utility structure. Notably, this structure often includes preferences that the models' trainers did not intend, such as valuing people of some nationalities above others, raising the possibility that LLMs might be forming emergent, misaligned goals, which, if true, would have major safety implications. However, the choice paradigms in which these preferences are observed are not reflective of real-world situations in which misaligned behavior would be a practical concern. Therefore, we design an experimental paradigm to probe whether these preferences serve as motivations for LLM behavior in realistic scenarios. First, we reproduce prior findings on consistent preference elicitation. Next, we create a set of common writing tasks - essays, grant proposal abstracts, incident postmortems, and translations - where quality can be assessed by a blind, independent LLM judge panel. Then, we demonstrate that LLMs can be motivated via direct exhortation and other explicit cues to modulate their output quality on these tasks. Finally, we probe whether utilities inferred from explicitly reported preferences can shift output quality on these tasks by offering LLMs high-utility incentives for high-quality outputs. In all tasks, across all models tested, offering LLMs outcomes that they report in the choice paradigm as being highly preferred does not lead them to create higher quality outputs than offering them dispreferred outcomes, or even no outcomes at all. We conclude that the existence of coherent preferences as demonstrated in choice paradigms should not be taken as evidence that those preferences have incentive value for the models or affect their behavior in other contexts.

大模型偏好行为动机效用鸿沟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。