arXiv:2410.01529cs.ROcs.CV2024-10ICRA被引 8

用单模态数据教会机器人理解多模态任务指令

Robo-MUTUAL: Robotic Multimodal Task Specification via Unimodal Learning

  • 通过海量单模态数据预训练多模态编码器,实现跨模态对齐
  • 引入两种扰动操作缩小模态差距,使不同模态表示可互换
  • 在130多个任务上验证,实现在数据稀缺下的强泛化能力

多模态任务描述对提升机器人性能至关重要,其中跨模态对齐使机器人能整体理解复杂指令。直接标注配对的多模态数据用于模型训练不切实际,因多模态数据稀疏。本研究证明,利用真实数据中丰富的单模态指令,可有效教会机器人学习多模态任务规范。首先,通过大量非领域数据预训练机器人多模态编码器,赋予其强跨模态对齐能力;随后,采用两种“坍缩与破坏”操作,进一步弥合学习到的多模态表征中的剩余模态差异。该方法将相同任务目标的不同模态表示投影为可互换形式,从而在对齐良好的多模态潜在空间中实现精准机器人操作。在超过130个任务、4000次评估中,于仿真LIBERO基准和真实机器人平台上的测试表明,所提框架显著克服了机器人学习中的数据限制,展现出卓越性能。

原文摘要 · Abstract (English)

Multimodal task specification is essential for enhanced robotic performance, where \textit{Cross-modality Alignment} enables the robot to holistically understand complex task instructions. Directly annotating multimodal instructions for model training proves impractical, due to the sparsity of paired multimodal data. In this study, we demonstrate that by leveraging unimodal instructions abundant in real data, we can effectively teach robots to learn multimodal task specifications. First, we endow the robot with strong \textit{Cross-modality Alignment} capabilities, by pretraining a robotic multimodal encoder using extensive out-of-domain data. Then, we employ two Collapse and Corrupt operations to further bridge the remaining modality gap in the learned multimodal representation. This approach projects different modalities of identical task goal as interchangeable representations, thus enabling accurate robotic operations within a well-aligned multimodal latent space. Evaluation across more than 130 tasks and 4000 evaluations on both simulated LIBERO benchmark and real robot platforms showcases the superior capabilities of our proposed framework, demonstrating significant advantage in overcoming data constraints in robotic learning. Website: zh1hao.wang/Robo_MUTUAL

机器人多模态单模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。