arXiv:2606.20624cs.AIcs.CL2026-06

大模型推理时存在理性缺失,即使对齐也未必最优。

In LLM Reasoning, there is Irrationality on top of Value Misalignment

  • 提出理性价值风险概念,量化推理策略与最优解的差距
  • 实验证明多数主流模型普遍存在理性缺陷,且无法完全避免
  • 自一致性与长思维链可提升理性,但收益递减

尽管大语言模型在对齐目标价值函数方面取得显著进展,我们指出,即便模型在(后)训练中已良好对齐,其推理过程仍可能无法最大化对齐后的价值。本文数学形式化该差距为理性价值风险:模型部署推理策略与理论最优策略(沿最陡方向最大化效用)之间的效用差异。理性价值风险的估计误差进一步分解为三部分:受限提示、受限响应与不完美验证器。实验覆盖Llama-3.1、Qwen-2.5、Tülu-3系列(7B-72B)、GPT-5.2、GPT-5.5和DeepSeek-V4,基准包括UltraFeedback、AlpacaEval、GSM8K、MATH、HumanEval和MathArena。结果表明:(1) 理性价值风险广泛存在;(2) 价值对齐虽能缓解,但无法消除;(3) 自一致性可提升理性;(4) 更长的思维链有助于理性,但边际收益递减。代码开源于https://github.com/EVIEHub/LLM-Rationality。

原文摘要 · Abstract (English)

Significant progress has been made in aligning LLMs with target value functions. We argue that, even when an LLM has been well aligned in (post-)training, it may still fail to maximise the aligned value in reasoning. We mathematically formalise this gap as rational value risk: the utility discrepancy between a model's deployed reasoning strategy and its rational counterpart whose responses maximise utility in the steepest direction. The estimation error of rational value risk is further decomposed into three components from bounded prompts, bounded responses, and imperfect verifiers. Extensive experiments are conducted, covering models Llama-3.1, Qwen-2.5, Tülu-3 families (7B-72B), GPT-5.2, GPT-5.5, and DeepSeek-V4, and benchmarks UltraFeedback, AlpacaEval, GSM8K, MATH, HumanEval, and MathArena. The results validate that (1) rational value risk is widespread; (2) value alignment can reduce, but cannot avoid, it; (3) self-consistency can improve rationality; and (4) a longer chain of thought improves rationality but with diminishing returns. The code is at https://github.com/EVIEHub/LLM-Rationality

大模型推理理性缺陷价值对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。