实证研究强化学习对大模型智能体泛化能力的影响
Can RL Improve Generalization of LLM Agents? An Empirical Study
- 通过三类泛化实验评估强化微调的跨任务、跨环境表现
- 在相同环境中任务难度变化时泛化良好,但跨环境迁移能力弱
- 顺序训练能提升下游性能且遗忘少,多环境混合训练更平衡
强化微调(RFT)在基于环境反馈训练大模型智能体进行多轮决策方面展现出潜力。然而,现有评估大多局限于同一环境内:训练与测试在同一环境甚至相同任务上进行。在真实部署中,智能体可能面临未见过的环境,其背景知识、观测空间和动作接口均不同。为刻画RFT在这些变化下的泛化特性,我们从三个维度开展系统性研究:(1)同一环境内任务难度的泛化;(2)跨环境迁移到未知环境;(3)序列式多环境训练以量化迁移与遗忘。结果表明,RFT在环境内任务难度变化下表现良好,但在跨环境迁移中较弱,且与语义先验及观测/动作接口变化相关。相比之下,序列训练带来显著下游收益且上游遗忘极少,跨环境混合训练则提升了整体平衡性。我们进一步提供详细分析与深入洞见,期望助力社区构建可部署的通用大模型智能体。
原文摘要 · Abstract (English)
Reinforcement fine-tuning (RFT) has shown promise for training LLM agents to perform multi-turn decision-making based on environment feedback. However, most existing evaluations remain largely in-domain: training and testing are conducted in the same environment or even on the same tasks. In real-world deployment, agents may operate in unseen environments with different background knowledge, observation spaces, and action interfaces. To characterize the generalization profile of RFT under such shifts, we conduct a systematic study along three axes: (1) within-environment generalization across task difficulty, (2) cross-environment transfer to unseen environments, and (3) sequential multi-environment training to quantify transfer and forgetting. Our results show that RFT generalizes well across task difficulty within an environment, but exhibits weaker transfer to unseen environments, which correlates with shifts in both semantic priors and observation/action interfaces. In contrast, sequential training yields promising downstream gains with minimal upstream forgetting, and mixture training across environments improves the overall balance. We further provide detailed analyses and deeper insights, and hope our work helps the community develop and deploy generalizable LLM agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。