arXiv:2605.27954cs.LG2026-05

发现智能体强化学习中熵的周期性爆发现象,并提出新方法稳定训练。

Cyclical Entropy Eruption: Entropy Dynamics in Agent Reinforcement Learning

论文配图:Cyclical Entropy Eruption: Entropy Dynamics in Agent Reinforcement Learning
图 1 · 摘自论文原文
  • 揭示智能体强化学习中熵周期性剧烈波动的新现象。
  • 提出SEAL辅助损失,有效降低错误轨迹的熵值,提升训练稳定性。
  • 适用于大模型智能体训练,尤其适合需长期推理的任务场景。

基于大语言模型的智能体通过目标推理、调用工具和与外部环境交互来解决现实任务,强化学习为此提供了自然框架。尽管近期方法在多个领域取得显著成果,但智能体强化学习的训练动态仍不清晰,限制了对不稳定性的诊断及更高效训练算法的设计。本文首次揭示了一种此前未被充分研究的现象——周期性熵爆发:与单轮推理强化学习中熵持续下降不同,智能体训练呈现熵值反复剧烈上升后逐渐回落的循环特征。我们将其分解为三个阶段,从理论与实证角度分析其机制。进一步发现,在熵爆发期间形成的重复句子、幻觉等退化模式会跨周期累积。受此启发,提出轻量级辅助损失 SEAL(分离增强智能体学习),通过在表征空间中分离正确与错误轨迹,直接针对熵爆发的根本原因。在多个基准测试、模型和强化学习算法上的实验表明,SEAL能显著稳定训练过程,并带来更强的下游智能体性能。

原文摘要 · Abstract (English)

Agentic large language models are increasingly used to solve real-world tasks by reasoning over goals, invoking tools, and interacting with external environments. Reinforcement learning provides a natural framework for improving these behaviors, and recent agent RL methods have achieved strong results across domains. However, the training dynamics of agent RL remain poorly understood, limiting our ability to diagnose instabilities and design more effective training algorithms. In this work, we identify a previously underexplored phenomenon in agent RL, which we term cyclical entropy eruption. Unlike single-turn reasoning RL, where entropy typically collapses and stays low, agent RL training exhibits unique recurring cycles of sharp entropy eruption and gradual subsidence. We decompose this dynamic into three phases and provide theoretical and empirical analyses of each, explaining the mechanisms underlying its cyclical oscillation. We further show that degenerate patterns such as sentence duplication and hallucination, once acquired during eruption, can persist and accumulate across cycles. Motivated by these findings, we propose SEAL (Separation-Enhanced Agent Learning), a lightweight auxiliary loss that separates correct and incorrect trajectories in representation space, directly targeting the root cause of entropy eruption. Experiments across multiple benchmarks, models, and RL algorithms demonstrate that SEAL stabilizes training and yields stronger downstream agent performance.

强化学习大模型智能体训练稳定熵动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。