arXiv:2505.10105cs.ROcs.AI2025-05被引 5

提出统一3D多模态表示EmbodiedMAE,提升机器人操作的感知与决策能力。

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation

  • 通过随机掩码和跨模态融合,统一学习RGB、深度图与点云信息。
  • 在70个仿真和20个真实任务中,性能超越现有视觉基础模型。
  • 适合需要精准空间感知的桌面操作场景,尤其对大模型有良好扩展性。

本文提出EmbodiedMAE,一种用于机器人操作的统一3D多模态表示方法。当前方法存在训练数据与实际操作间显著领域差异,且缺乏有效整合3D信息的模型架构。为此,我们基于DROID数据集增强高质量深度图与点云,构建了适用于3D具身视觉研究的DROID-3D数据集。随后设计EmbodiedMAE,一种多模态掩码自编码器,通过随机掩码与跨模态融合,同步学习RGB、深度图与点云的表征。在DROID-3D上训练后,EmbodiedMAE在70个仿真任务与20个真实机器人操作任务(两个平台)中均展现出优于当前最优视觉基础模型(VFMs)的训练效率与最终性能。模型表现出良好的规模扩展性,能有效支持从3D输入出发的策略学习。实验表明,EmbodiedMAE是具身智能系统中可靠的统一3D多模态视觉基础模型,特别适用于对空间感知要求高的精确桌面操作场景。

原文摘要 · Abstract (English)

We present EmbodiedMAE, a unified 3D multi-modal representation for robot manipulation. Current approaches suffer from significant domain gaps between training datasets and robot manipulation tasks, while also lacking model architectures that can effectively incorporate 3D information. To overcome these limitations, we enhance the DROID dataset with high-quality depth maps and point clouds, constructing DROID-3D as a valuable supplement for 3D embodied vision research. Then we develop EmbodiedMAE, a multi-modal masked autoencoder that simultaneously learns representations across RGB, depth, and point cloud modalities through stochastic masking and cross-modal fusion. Trained on DROID-3D, EmbodiedMAE consistently outperforms state-of-the-art vision foundation models (VFMs) in both training efficiency and final performance across 70 simulation tasks and 20 real-world robot manipulation tasks on two robot platforms. The model exhibits strong scaling behavior with size and promotes effective policy learning from 3D inputs. Experimental results establish EmbodiedMAE as a reliable unified 3D multi-modal VFM for embodied AI systems, particularly in precise tabletop manipulation settings where spatial perception is critical.

3D视觉多模态机器人操作自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。