arXiv:2412.07215cs.ROcs.MM2024-12ICCV被引 18

首个通用机器人操作大模型,实现多数据集统一评估

RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation

  • 融合视觉、深度与相机参数,提升3D空间感知能力
  • 在CALVIN数据集上平均序列长度从1.7提升至3.5,超越专家模型
  • 支持跨实体泛化,适合需要统一评估的机器人研究者

近期,机器人领域通过引入更大模型和大规模数据集取得了显著进展。然而,在3D空间交互建模和数据采集成本方面仍面临挑战。为此,我们提出多模态机器人操作模型RoboTron-Mani及综合性数据集RoboData。RoboTron-Mani一方面通过相机参数和占用监督增强3D感知;另一方面基于OpenFlamingo引入模态隔离掩码与多模态解码模块,提升模态融合与细粒度理解。RoboData整合多个公开数据集,首次实现多视角图像、相机参数、深度图、动作与空间对齐的统一,支持多样化机器人数据的综合学习,并提供完整评估体系。在RoboData上训练的RoboTron-Mani是首个超越专家模型的通用策略,可同时在多个数据集上评估所有任务,而非局限于特定数据或任务。具体而言,其在CALVIN数据集上的平均序列长度由1.7提升至3.5,实现跨实体泛化,并在模拟与真实世界数据集上均取得最先进性能。

原文摘要 · Abstract (English)

Recently, robotics has advanced significantly through the integration of larger models and large-scale datasets. However, challenges remain in applying these models to 3D spatial interactions and managing data collection costs. To address these issues, we propose the multimodal robotic manipulation model RoboTron-Mani and the comprehensive dataset RoboData. RoboTron-Mani, on one hand, enhances 3D perception through camera parameters and occupancy supervision. On the other hand, it further incorporates Modality-Isolation-Mask and multimodal decoder blocks based on OpenFlamingo, improving modality fusion and fine-grained perception. RoboData integrats several publicly-available datasets, achieving the first fusion of multi-view images, camera parameters, depth maps, actions, and space alignment, which facilitates comprehensive learning from diverse robotic datasets and offers one complete evaluation system. Trained on RoboData, RoboTron-Mani is the first generalist policy that surpasses expert models, enabling simultaneous evaluation of all tasks across multiple datasets, rather than being limited to specific data or task selections. Specifically, RoboTron-Mani boosts manipulation performance by increasing the average sequence length on CALVIN from 1.7 to 3.5, enabling cross-embodiment generalization, and achieving state-of-the-art results on both simulated and real-world datasets.

机器人操作多模态通用策略数据融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。