arXiv:2510.21339cs.CLcs.IT2025-10

单轮训练比多轮反馈训练更有效,能更好提升大模型推理能力。

Multi-turn Training with Basic Human Feedback Helps Little on LLM Reasoning

  • 用单轮强化学习训练模型,效果优于多轮人类反馈训练。
  • 多轮训练使单轮推理性能下降约15%以上,影响模型可靠性。
  • 适合追求稳定推理的场景,如需要快速决策的任务。

大型语言模型(LLMs)的推理能力通常通过单轮强化学习进行训练,但实际应用中常涉及多轮人类反馈交互,导致训练与部署条件不一致。本文研究多轮人类反馈训练是否对推理任务必要。对比传统单轮训练与三种多轮策略,发现单轮训练模型在单轮和多轮评估中均表现良好,而多轮训练模型在单轮推理中性能显著下降。结果表明,对于信息完整的任务,稳健的单轮训练仍更有效可靠,多轮基础反馈训练带来的收益有限,甚至会削弱推理能力。

原文摘要 · Abstract (English)

The reasoning capabilities of Large Language Models (LLMs) are typically developed through the single-turn reinforcement learning, whereas real-world applications often involve multi-turn interactions with human feedback, leading to a potential mismatch between training and deployment conditions. In this work, we study whether multi-turn training with human feedback is necessary for reasoning tasks. We compare conventional single-turn training with three multi-turn strategies and reach contrary conclusions to previous research. We find that models trained in a single-turn setting generalize effectively to both single- and multi-turn evaluations, while models trained with multi-turn strategies exhibit a significant degradation in single-turn reasoning performance. These results suggest that for tasks with complete information, robust single-turn training remains more effective and reliable, as multi-turn training with basic feedback provides limited benefits and can even degrade reasoning capabilities.

大模型推理强化学习训练策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。