用物体为中心的技能表示,让机器人更高效地学习和泛化操作任务。
Generalizable Hierarchical Skill Learning via Object-Centric Representation
- 以物体为中心分解示范,构建可迁移的技能基元
- 3次演示即达基线30倍数据效果,未见场景下提升15.5%
- 适合需要少样本、强泛化的机器人操作场景
我们提出通用分层技能学习(GSL),一种新型分层策略学习框架,显著提升机器人操作中的策略泛化能力与样本效率。GSL的核心思想是使用物体为中心的技能作为视觉-语言模型与低层视觉-运动策略之间的接口。具体而言,GSL利用基础模型将示范分解为可迁移且物体规范化的技能基元,确保在物体坐标系下高效学习低层技能。测试时,高层代理预测的技能-物体对输入低层模块,推断出的规范动作被映射回世界坐标系执行。这种结构化且灵活的设计在未见过的空间布局、物体外观和任务组合上均实现显著提升。仿真中,仅需每任务3次示范的GSL,在未见任务上表现优于基线(30倍数据训练)15.5%;真实世界实验中,也超越了使用10倍更多数据训练的基线。
原文摘要 · Abstract (English)
We present Generalizable Hierarchical Skill Learning (GSL), a novel framework for hierarchical policy learning that significantly improves policy generalization and sample efficiency in robot manipulation. One core idea of GSL is to use object-centric skills as an interface that bridges the high-level vision-language model and the low-level visual-motor policy. Specifically, GSL decomposes demonstrations into transferable and object-canonicalized skill primitives using foundation models, ensuring efficient low-level skill learning in the object frame. At test time, the skill-object pairs predicted by the high-level agent are fed to the low-level module, where the inferred canonical actions are mapped back to the world frame for execution. This structured yet flexible design leads to substantial improvements in sample efficiency and generalization of our method across unseen spatial arrangements, object appearances, and task compositions. In simulation, GSL trained with only 3 demonstrations per task outperforms baselines trained with 30 times more data by 15.5 percent on unseen tasks. In real-world experiments, GSL also surpasses the baseline trained with 10 times more data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。