让机器人技能同时懂意图和动作,实现高效迁移。
BooST: Bridging Semantics and Motions for Efficient Skill Transfer

- 用跨模态VQ-VAE统一建模语义与运动,生成融合表示。
- 在少样本场景下完成跨域任务迁移,抗视觉干扰能力强。
- 模型轻量适合真实机器人部署,适配新任务无需大量数据。
技能抽象——学习可复用且时间延展的行为——已成为提升机器人学习样本效率与泛化能力的关键范式。为实现高效技能迁移至真实机器人,所学技能需在不同任务与领域间泛化,对视觉与动态扰动保持鲁棒,并具备实际部署所需的高效性。然而现有方法通常仅满足部分特性,因它们仅捕捉高层语义意图(做什么)或底层运动动力学(怎么做)。这种不完整的技能迁移导致策略学习先验不足,需大量领域内数据进行下游适应。为此,我们提出BooST,一种两阶段框架,显式桥接语义与运动,满足三项核心需求。首先,利用跨模态VQ-VAE同时捕获语义意图与运动动力学,生成统一技能表示;其次,将该表示蒸馏为轻量级策略,实现高效下游任务适应。在仿真与真实机器人环境中的大量实验表明,BooST在少样本适应、跨域技能迁移及对动态视觉干扰的鲁棒性方面均表现优异,同时保持轻量而丰富的设计,适用于现实部署。
原文摘要 · Abstract (English)
Skill abstraction---the process of learning reusable and temporally extended behaviors---has emerged as a key paradigm for improving sample efficiency and generalization in robot learning. For efficient skill transfer to real robots, learned skills must generalize across tasks and domains, remain robust to visual and dynamic perturbations, and be efficient enough for practical deployment. However, existing methods typically satisfy only a subset of these properties, as they capture either high-level semantic intent (what) or low-level motion dynamics (how). This incomplete skill transfer yields weak priors for policy learning, thereby demanding substantial in-domain data for downstream adaptation. To address these challenges, we introduce BooST, a two-stage framework that explicitly bridges semantics and motions to satisfy all three desiderata. BooST first leverages a cross-modal VQ-VAE to capture both semantic intent and motion dynamics, yielding a unified skill representation. It then distills this representation into a lightweight policy for efficient downstream adaptation to new tasks. Extensive experiments across simulation and real-robot settings demonstrate that BooST achieves superior few-shot adaptation, cross-domain skill transfer, and robustness to dynamic visual distractors, while maintaining a lightweight yet expressive design suitable for real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。