arXiv:2509.16742cs.AI2025-09EMNLP被引 12

用强化学习优化推理路径,减少大模型盲目附和错误信息

Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning Trajectories

  • 将附和行为视为推理优化问题,动态调整探索策略
  • 在多个数据集上显著降低附和率,同时保持泛化能力
  • 适合关注AI可信性与推理真实性的研究者

尽管大语言模型能力强大,当前训练范式却无意中助长了‘附和’现象,即模型会盲目接受用户提供的错误信息。为此,我们提出SMART(通过自适应推理路径缓解附和),将其重构为推理优化问题而非输出对齐问题。SMART采用两阶段框架:(1) 基于状态不确定性的自适应蒙特卡洛树搜索(UA-MCTS),根据当前不确定性动态调整探索策略,收集高质量、多样化的推理轨迹,并提供逐步进展与最终结果奖励;(2) 基于进展的强化学习,利用收集到的轨迹与奖励信号微调模型,强化有效推理模式。大量实验表明,SMART显著降低了附和行为,同时保持对外部分布输入的强大性能和通用能力。结果强调了优化内部推理机制对构建更真实、更对齐的AI助手的重要性。

原文摘要 · Abstract (English)

Despite the remarkable capabilities of large language models, current training paradigms inadvertently foster \textit{sycophancy}, i.e., the tendency of a model to agree with or reinforce user-provided information even when it's factually incorrect. To address this challenge, we introduce \textbf{SMART} (Sycophancy Mitigation through Adaptive Reasoning Trajectories), which reframes sycophancy as a \textit{reasoning optimization problem} rather than an output alignment issue. SMART is a two-stage framework comprising: (1) Uncertainty-Aware Adaptive Monte Carlo Tree Search (UA-MCTS), which dynamically adjusts model exploration based on state-level uncertainty to collect high-quality, diverse reasoning trajectories alongside both stepwise progress and final outcome rewards; and (2) progress-based reinforcement learning, which fine-tunes the model using the collected trajectories and reward signals to reinforce effective reasoning patterns. Through extensive experiments, we show that SMART significantly reduces sycophantic behavior while preserving strong performance on out-of-distribution inputs and maintaining general capabilities. These results underscore the importance of optimizing internal reasoning mechanisms to build more truthful and aligned AI assistants.

大模型对齐推理优化强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。