arXiv:2604.11611cs.CLcs.LG2026-04

用自评估生成奖励,让小模型像大模型一样学会复杂任务

Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation

论文配图:Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation
图 1 · 摘自论文原文
  • 通过事后生成自我评价提供密集奖励信号
  • 70亿参数模型在无专家监督下达到GPT-4o水平表现
  • 理论证明奖励校准可提升学习效率,适合自主智能体研究

为解决基于大语言模型(LLM)的强化学习中奖励稀疏问题,我们提出互信息自评估(MISE)机制,利用事后生成的自评估作为密集奖励信号,并同步将其与环境反馈进行校准。实验表明,MISE使代理能够仅依靠内部密集奖励和稀疏外部信号实现自主学习。理论上,本工作首次为生成式自奖励范式提供了形式化基础,证明使用事后自评估奖励等价于最小化互信息与策略和代理奖励策略间KL散度之和的目标函数。该理论洞察指导并验证了我们的校准步骤,主动将奖励对齐最优策略。大量实验显示,MISE显著优于强基线,在无需专家监督的情况下,使约70亿参数的开源大模型在验证集上达到GPT-4o性能。

原文摘要 · Abstract (English)

To overcome the sparse reward challenge in reinforcement learning (RL) for agents based on large language models (LLMs), we propose Mutual Information Self-Evaluation (MISE), an RL paradigm that utilizes hindsight generative self-evaluation as dense reward signals while simultaneously calibrating them against the environmental feedbacks. Empirically, MISE enables an agent to learn autonomously from dense internal rewards supplementing sparse extrinsic signals. Theoretically, our work provides the first formal foundation for the paradigm of generative self-rewarding. We prove that utilizing hindsight self-evaluation rewards is equivalent to minimizing an objective that combines mutual information with a KL divergence term between the policy and a proxy reward policy. This theoretical insight then informs and justifies our calibration step, which actively aligns these rewards with the optimal policy. Extensive experiments show that MISE outperforms strong baselines, enabling open-source LLMs about 7B parameters to achieve performance comparable to GPT-4o on validation without expert supervision.

强化学习自评估语言模型奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。