arXiv:2601.22900cs.AI2026-01被引 1

用多轮语言反馈提升强化学习的推理能力,让模型从错误中学会改进。

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop

  • 通过多轮语言反馈引导失败样本的重生成,实现渐进式优化
  • 在OpenR1-Math上超越监督学习与基线强化学习方法
  • 适合需要复杂推理与自我修正的AI系统开发者

基于可验证奖励的强化学习(RLVR)广泛用于提升跨领域的推理能力,但仅依赖结果的标量奖励往往稀疏且缺乏信息。尤其在失败样本中,标量奖励仅表明答案错误,却无法解释推理过程的断裂点。本文提出MulFeRL(多轮反馈引导强化学习),一种事件触发的多轮RLVR框架,结合反馈驱动的进度诱导、验证确认后的进度信用分配,以及结构化反馈注入机制,将语言反馈转化为可训练的学习信号。在采样的OpenR1-Math数据集上训练后,MulFeRL在域内表现优于监督学习、自蒸馏及传统RLVR基线,并展现出出色的跨域泛化能力。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable Rewards (RLVR) is widely used to improve reasoning across domains, but outcome-only scalar rewards are often sparse and uninformative. This limitation is especially severe for failed samples, where scalar rewards indicate only that a solution is incorrect without explaining why the reasoning breaks down. In this paper, we leverage richer verbal feedback to guide RLVR on failed samples and convert feedback-induced progress into trainable learning signals. We propose MulFeRL (Multi-turn Feedback-guided Reinforcement Learning), a multi-turn, event-triggered RLVR framework that combines progress induction for feedback-guided regeneration of failed samples, progress credit assignment for learning from verifier-confirmed progress, and structured feedback injection for integrating feedback into the model's reasoning process. Trained on sampled OpenR1-Math, MulFeRL outperforms supervised, self-distillation-based, and RLVR baselines in-domain, while also showing strong out-of-domain generalization.

强化学习语言反馈推理增强多轮交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。