arXiv:2509.20612cs.LG2025-09NeurIPS被引 2

让智能体新学的技能自动适配旧策略,无需重训

Policy Compatible Skill Incremental Learning via Lazy Learning Interface

  • 用双侧懒惰学习动态对齐策略子任务与技能空间
  • 新技能提升后,下游策略性能不降反升
  • 适合长期迭代、需保持策略兼容性的机器人应用

技能增量学习(SIL)使具身智能体通过环境交互或新增数据持续扩展和优化技能集,从而高效构建可复用的分层策略。然而,随着技能库演进,原有基于技能的策略可能因不兼容而失效,限制其泛化能力。本文提出SIL-C框架,通过双侧懒惰学习映射机制,动态对齐策略所依赖的子任务空间与智能体行为解码的技能空间。该方法使每个子任务可通过轨迹分布相似性匹配最优技能执行。在多种SIL场景下的实验表明,SIL-C在保持技能与策略兼容性的同时,确保了整个学习过程的效率。

原文摘要 · Abstract (English)

Skill Incremental Learning (SIL) is the process by which an embodied agent expands and refines its skill set over time by leveraging experience gained through interaction with its environment or by the integration of additional data. SIL facilitates efficient acquisition of hierarchical policies grounded in reusable skills for downstream tasks. However, as the skill repertoire evolves, it can disrupt compatibility with existing skill-based policies, limiting their reusability and generalization. In this work, we propose SIL-C, a novel framework that ensures skill-policy compatibility, allowing improvements in incrementally learned skills to enhance the performance of downstream policies without requiring policy re-training or structural adaptation. SIL-C employs a bilateral lazy learning-based mapping technique to dynamically align the subtask space referenced by policies with the skill space decoded into agent behaviors. This enables each subtask, derived from the policy's decomposition of a complex task, to be executed by selecting an appropriate skill based on trajectory distribution similarity. We evaluate SIL-C across diverse SIL scenarios and demonstrate that it maintains compatibility between evolving skills and downstream policies while ensuring efficiency throughout the learning process.

增量学习策略兼容懒惰学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。