让大模型自主掌握技能,告别依赖外部提示。
SKILLC: Learning Autonomous Skill Internalization in LLM Agents via Contrastive Credit Assignment

- 用对比奖励机制区分技能辅助与独立成功,指导内部化学习。
- 在ALFWorld和WebShop上超越现有内化方法5.5%和4.4%。
- 适合研究智能体自主性、技能内化的研究人员。
结构化技能提示能提升长程智能体强化学习中的探索效率。现有技能内化方法仅用技能有用性对比来控制课程,未改变策略更新,无法区分依赖技能与自主成功。本文提出SkillC框架,基于对比技能信用分配(CSCA),将该对比转化为直接学习信号。SkillC在同一策略更新中采样同一任务类型下的带技能与无技能轨迹,通过双流优势估计器将任务级对比注入优化,保持全局排序的同时对无技能成功进行单向修正。平滑的验证级信号则动态调节归因强度、轨迹分配及单调活跃集剪枝。在ALFWorld和WebShop上的实验表明,无需运行时技能访问,SkillC分别比最强先期内化基线高出5.5%和4.4%,且仍与技能增强型方法相当。
原文摘要 · Abstract (English)
Structured skill prompts improve exploration in long-horizon agentic reinforcement learning (RL). Skill-augmented RL methods retain external skills at inference, while skill-internalization RL methods withdraw them during training to enable autonomous performance. However, existing internalization approaches only use skill-helpfulness contrast for curriculum control, leaving the policy update unchanged and unable to distinguish skill-dependent from autonomous success. We propose SkillC, a framework based on Contrastive Skill Credit Assignment (CSCA) that converts this contrast into a direct learning signal for internalization. \textsc{SkillC} samples paired skill-injected and skill-free rollouts for tasks from active skill types within the same policy update, and injects their task-level contrast into optimization via a dual-stream advantage estimator that preserves global ranking while applying a one-sided correction toward skill-free success. A smoothed validation-level signal further drives an adaptive curriculum over attribution strength, rollout allocation, and monotonic active-set pruning. Experiments on ALFWorld and WebShop show that, without runtime skill access, SkillC surpasses the strongest prior skill-internalization RL baseline by 5.5\% and 4.4\%, respectively, while remaining competitive with skill-augmented RL methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。