提出稳定智能体强化学习框架,解决训练崩溃难题
ARLArena: A Unified Framework for Stable Agentic Reinforcement Learning
- 构建标准化测试平台,分解策略梯度为四大设计维度
- 提出SAMPO方法,在多任务中实现稳定训练与强性能
- 提供可复现的智能体训练指南,适合构建大模型代理系统
智能体强化学习(ARL)作为解决复杂多步交互任务的新兴范式备受关注。尽管初期成果令人鼓舞,但其训练极不稳定,常导致训练崩溃,限制了在更大环境和更长交互周期中的扩展性,并阻碍算法设计的系统性探索。本文提出ARLArena,一个稳定训练配方与系统分析框架,可在可控且可复现的环境中研究训练稳定性。ARLArena首先构建清晰标准化的测试平台,将策略梯度分解为四个核心设计维度,并评估各维度的性能与稳定性。通过细粒度分析,提炼出统一的ARL视角,提出SAMPO——一种旨在缓解ARL主要不稳定性来源的稳定智能体策略优化方法。实证表明,SAMPO在多种智能体任务中均实现持续稳定训练与优异表现。本研究为ARL提供了统一的策略梯度视角,并为构建稳定可复现的大语言模型代理训练流水线提供了实用指导。
原文摘要 · Abstract (English)
Agentic reinforcement learning (ARL) has rapidly gained attention as a promising paradigm for training agents to solve complex, multi-step interactive tasks. Despite encouraging early results, ARL remains highly unstable, often leading to training collapse. This instability limits scalability to larger environments and longer interaction horizons, and constrains systematic exploration of algorithmic design choices. In this paper, we first propose ARLArena, a stable training recipe and systematic analysis framework that examines training stability in a controlled and reproducible setting. ARLArena first constructs a clean and standardized testbed. Then, we decompose policy gradient into four core design dimensions and assess the performance and stability of each dimension. Through this fine-grained analysis, we distill a unified perspective on ARL and propose SAMPO, a stable agentic policy optimization method designed to mitigate the dominant sources of instability in ARL. Empirically, SAMPO achieves consistently stable training and strong performance across diverse agentic tasks. Overall, this study provides a unifying policy gradient perspective for ARL and offers practical guidance for building stable and reproducible LLM-based agent training pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。