arXiv:2605.00425cs.AI2026-05被引 3

提出无需监督的自适应熵调节方法,提升大模型智能体多轮强化学习的信用分配效果。

AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning

  • 基于响应级熵动态调节,实现无需中间监督的探索与利用平衡
  • 在多个基准上显著提升基线性能,软件工程任务增益达1.4%
  • 适合追求高泛化能力的多轮决策类任务研究者

强化学习显著提升了大语言模型智能体与环境交互并完成多轮任务的能力。然而,有效的智能体强化学习仍具挑战:稀疏的结果奖励难以对长轨迹中各步骤进行有效信用分配。现有方法常引入密集的中间监督(如过程奖励模型或辅助自监督信号),增加调优复杂度且可能限制跨任务泛化。本文提出AEM,一种无监督的信用分配方法,通过自适应调节强化学习训练中的熵动态,优化探索与利用权衡。由于智能体的环境影响通常由完整响应而非单个标记决定,我们分析将熵动态从词元层级提升至响应层级,使不确定性估计与智能体有效动作粒度对齐,降低对词元级采样噪声的敏感性。进一步发现,在自然梯度更新下熵漂移受采样响应优势与其相对意外度的相互作用支配。基于此,AEM推导出实用的响应级不确定性代理,并用于重标优势,利用正负样本间动态平衡,自然实现从探索到利用的过渡。在ALFWorld、WebShop和SWE-bench-Verified上的大量实验表明,AEM持续改进强基线,集成至先进软件工程强化学习框架时获得+1.4%的性能提升。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has substantially improved the ability of large language model (LLM) agents to interact with environments and solve multi-turn tasks. However, effective agentic RL remains challenging: sparse outcome-only rewards provide limited guidance for assigning credit to individual steps within long interaction trajectories. Existing approaches often introduce dense intermediate supervision, such as process reward models or auxiliary self-supervised signals, which increases supervision and tuning complexity and may limit generalization across tasks and domains. We present AEM, a supervision-free credit assignment method that adaptively modulates entropy dynamics during RL training to improve the exploration-exploitation trade-off. Since in agentic RL the environment is typically affected by a complete response, rather than an individual token, our analysis lifts entropy dynamics from the token level to the response level, aligning uncertainty estimation with the effective action granularity of LLM agents and reducing sensitivity to token-level sampling noise. We further show that entropy drift under natural-gradient updates is governed by the interaction between the sampled-response advantage and its relative surprisal. Motivated by this result, AEM derives a practical response-level uncertainty proxy and uses it to rescale advantages, leveraging the evolving balance between positive and negative samples to naturally transition from exploration to exploitation. Extensive experiments on ALFWorld, WebShop, and SWE-bench-Verified with models ranging from 1.5B to 32B demonstrate that AEM consistently improves strong RL baselines, including a +1.4\% gain when integrated into a state-of-the-art software-engineering RL training framework.

强化学习大模型智能体信用分配自适应熵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。