arXiv:2510.24095cs.LGcs.AI2025-10NeurIPS被引 1

从专家示范中自动学习可调节的通用技能,提升新任务泛化能力。

Learning Parameterized Skills from Demonstrations

  • 联合学习元策略与参数化技能策略,实时选择技能及参数。
  • 在LIBERO和MetaWorld上超越多任务与基线方法,泛化性能显著。
  • 发现可解释技能,如可调抓取位置的物体抓取动作。

我们提出DEPS,一种端到端算法,用于从专家示范中发现参数化技能。该方法联合学习参数化技能策略与元策略,后者在每个时间步选择合适的离散技能和连续参数。通过结合时序变分推断与信息论正则化方法,解决潜在变量模型中常见的退化问题,确保学习到的技能具有时间延续性、语义意义且可适应。实验证明,从多任务专家示范中学习参数化技能能显著提升对未见任务的泛化能力。DEPS在LIBERO和MetaWorld基准测试中均优于多任务及技能学习基线方法。我们还展示了DEPS能发现可解释的参数化技能,例如其连续参数定义抓取位置的物体抓取技能。

原文摘要 · Abstract (English)

We present DEPS, an end-to-end algorithm for discovering parameterized skills from expert demonstrations. Our method learns parameterized skill policies jointly with a meta-policy that selects the appropriate discrete skill and continuous parameters at each timestep. Using a combination of temporal variational inference and information-theoretic regularization methods, we address the challenge of degeneracy common in latent variable models, ensuring that the learned skills are temporally extended, semantically meaningful, and adaptable. We empirically show that learning parameterized skills from multitask expert demonstrations significantly improves generalization to unseen tasks. Our method outperforms multitask as well as skill learning baselines on both LIBERO and MetaWorld benchmarks. We also demonstrate that DEPS discovers interpretable parameterized skills, such as an object grasping skill whose continuous arguments define the grasp location.

技能学习参数化强化学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。