让机器人通过动态感知学会快速适应新任务,少样本也能高效泛化。
Dynamics-Aware Meta-Imitation for Generalization to Unseen Robotic Manipulation

- 用元学习构建共享技能空间,结合视觉-运动轨迹捕捉动态规律。
- 在模拟和真实环境中,少样本微调后对未见任务的性能超越现有方法。
- 适合需要快速适应新操作任务的机器人系统研发人员使用。
模仿学习旨在从大量观测与示范中为机器人学习技能,但面临数据稀缺和环境泛化难题。现有方法多聚焦于域内任务的模仿,难以泛化到未见过的任务。为此,我们提出动力学感知的元模仿框架(DAMI)。通过元学习构建共享技能空间,使智能体能快速适应新任务。引入视觉-运动轨迹(VMT)模块以捕捉任务潜在空间中的复杂时空动态;提出无配对统一任务(U2T)块融合非结构化多模态观测;并设计任务条件特征调制(TCFM)机制,定制化调节低层3D特征。通过从随机完整参考示范中捕捉内在动态,框架学习任务本质逻辑而非记忆静态线索,确保有效泛化。大量仿真与真实世界实验表明,该方法在已见任务直接推理及未见任务少样本微调适应方面均优于当前最优基线。
原文摘要 · Abstract (English)
Imitation Learning aims to learn skills from extensive observations and demonstrations for robots, so it suffers from data scarcity and environment generalization. The existing methods predominantly focus on imitation from in-domain tasks and consequently struggle with generalization to unseen tasks. To bridge this generalization gap, we propose the \textbf{D}ynamics-\textbf{A}ware \textbf{M}eta-\textbf{I}mitation (DAMI) framework. By integrating meta-learning to construct a shared skill space, DAMI equips agents for rapid adaptation to novel tasks. We introduce the Visual-Motor Trajectory (VMT) module to capture complex spatio-temporal dynamics within the task latent space. Furthermore, we propose the Unpaired Unified Task (U2T) block to fuse unstructured multimodal observations. To coordinate these representations, we integrate a Task-Conditioned Feature Modulation (TCFM) mechanism customized for modulating low-level 3D features. By capturing intrinsic dynamics from a random complete reference demonstration, our framework learns the underlying task logic rather than memorizing static cues, ensuring effective generalization. Extensive experiments in both simulation and real-world settings demonstrate that our approach outperforms state-of-the-art baselines regarding direct inference on seen tasks and adaptation to unseen tasks via few-shot fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。