arXiv:2607.17946cs.LGcs.AI2026-07中稿 · ICML

通过几何视角优化推理路径,提升大模型在价值冲突下的决策稳定性。

A Geometric Perspective on Stabilizing Value Conflict Resolution

论文配图:A Geometric Perspective on Stabilizing Value Conflict Resolution
图 1 · 摘自论文原文
  • 用链式思维缓解强化学习中奖励压缩导致的优化不稳。
  • 在多个道德推理任务上,新设计的推理机制显著提升表现。
  • 适合关注大模型对齐与复杂价值判断的研究者参考。

大型语言模型在使用人类反馈强化学习(RLHF)训练时,常因压缩的标量奖励而难以应对价值冲突。本文从几何角度研究链式思维(CoT)的作用,发现其能有效平滑损失曲面中最陡峭方向,缓解传统标量奖励带来的优化不稳定性。通过下游基准测试验证,聚焦价值冲突的链式思维具备跨类型道德推理的泛化能力,展现出改善模型道德推理潜力。为此,我们提出一种新的价值冲突导向链式思维设计,进一步平滑损失曲面最陡方向,提升道德推理性能。结果表明,显式优化推理动态是提升大模型处理复杂价值冲突请求表现的有效途径,推动了大模型的多元对齐进展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) often struggle to navigate value conflicts when trained with the compressed scalar rewards of Reinforcement Learning from Human Feedback (RLHF). To address this challenge, we investigate how chain-of-thought (CoT) reasoning can help improve performance in this domain. Geometrically, we show that CoT correlates with further smoothing the model's loss landscape in its sharpest direction, helping resolve the optimization instability of traditional scalar rewards. We also demonstrate via relevant downstream benchmarks that value conflict-focused CoT may generalize to different kinds of moral reasoning, demonstrating that this CoT has the potential to be an effective mechanism for better moral reasoning. To capitalize on this potential, we create a new value conflict-focused CoT design that further smooths the sharpest direction of the loss landscape and increases moral reasoning performance. This finding shows that explicitly modifying and improving the design of reasoning dynamics offers a promising avenue for improving model performance on user requests with complex value conflicts, advancing pluralistic alignment in LLMs.

大模型对齐链式思维价值冲突

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。