提出Skill0.5框架,实现通用技能内化与特定技能调用的动态平衡。
Skill0.5: Joint Skill Internalization and Utilization for Out-of-Distribution Generalization in Agentic Reinforcement Learning

- 通过难度感知路由区分任务层级,分别采用内化与调用策略
- 在ALFWorld和WebShop上提升分布内/外泛化性能,优于基线方法
- 适合需要长期记忆与灵活执行的智能体系统研究者
赋予大语言模型显式技能已成为实现自主智能体解决复杂任务的有前景范式。智能体技能可自然划分为适用于广泛认知迁移的通用技能和用于动态执行的任务特定技能。然而,现有基于技能的强化学习方法通常在完全外部化(导致上下文开销过大)与完全内化(存在过拟合与知识冲突风险)之间做出僵化选择。为解决此困境,我们提出Skill0.5,一种新型智能体强化学习框架,通过将通用技能内化与任务特定技能调用相结合,明确区分技能处理方式。该框架由动态难度感知路由器驱动,将任务流分入不同掌握层级,并应用定制优化策略:对困难任务通过特权蒸馏内化通用技能以构建认知基础;对简单任务则采用诊断探测惩罚捷径,强制特定技能使用。在ALFWorld和WebShop上的实验表明,Skill0.5优于基于记忆和基于技能的强化学习基线,在分布内与分布外场景下均实现性能提升。
原文摘要 · Abstract (English)
Equipping large language models with explicit skills has emerged as a promising paradigm for enabling autonomous agents to solve complex tasks. Agent skills can be inherently divided into general skills for broad cognitive transfer and task-specific skills for dynamic execution. However, existing skill-based reinforcement learning (RL) methods typically force a rigid choice between full externalization, which incurs prohibitive context overhead, and full internalization, which risks overfitting and knowledge conflicts. To address this dilemma, we propose Skill0.5, a novel agentic RL framework that explicitly differentiates skill treatments by combining general skill internalization with task-specific skill utilization. Driven by a dynamic, difficulty-aware router, Skill0.5 streams tasks into distinct mastery tiers to apply tailored optimization strategies: it internalizes general skills via privileged distillation to build a cognitive foundation for hard tasks, while using diagnostic probing on easy tasks to penalize shortcuts and enforce specific skill utilization. Experiments on ALFWorld and WebShop demonstrate that Skill0.5 outperforms both memory-based and skill-based RL baselines, yielding performance improvements across both in-distribution and out-of-distribution scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。