arXiv:2502.07837cs.ROcs.LG2025-02被引 9

RoboBERT用两阶段训练实现高效多模态机器人操作,无需大量微调。

RoboBERT: An End-to-end Multimodal Robotic Manipulation Model

  • 分两阶段训练:先固定视觉编码器学稳定策略,再解冻模块注入多样指令。
  • 在CALVIN ABCD-D上达4.52的平均任务长度,优于现有方法。
  • 适合追求高效、轻量级多模态机器人系统的研究者和开发者。

具身智能融合视觉、语言与动作。然而,多数多模态机器人模型依赖大规模微调,成本高昂。为此,我们提出RoboBERT,一种基于新型两阶段训练范式的端到端多模态操作模型。第一阶段冻结大部分视觉编码器,仅用单一标准指令表述训练基于CNN的扩散策略,聚焦于稳定策略学习。第二阶段解冻所有模块,注入多样自然语言变体,快速对齐不同指令至已学策略,且不破坏性能。进一步采用系统性数据增强提升对视觉扰动的鲁棒性。不依赖辅助数据集,仅使用语言标注专家示范,RoboBERT在CALVIN ABCD-D基准上达到4.52的平均任务长度,在ABC-D上达3.79,刷新当前最优表现。6-DOF机械臂实机测试显示,其成功率高于相同数据训练的可比方法。结果表明,该数据增强强化的两阶段训练范式能为多模态机器人系统提供高效、可扩展且通用的性能。

原文摘要 · Abstract (English)

Embodied intelligence seamlessly integrates vision, language, and action.~However, most multimodal robotic models rely on massive fine-tuning, incurring high time and hardware costs.~To address this, we introduce RoboBERT, an end-to-end multimodal manipulation model built around a novel two-stage training paradigm.~In the first stage, we freeze most of the vision encoder and train with a single "standard" instruction phrasing, allowing the model to focus on stable policy learning via a CNN-based diffusion policy.~In the second stage, we unfreeze all modules and inject diverse natural language variants, rapidly aligning varied instructions to the already-learned policy without destabilizing performance.~We further employ systematic data augmentations to enhance robustness against visual perturbations.~Without relying on auxiliary datasets, RoboBERT achieves new state-of-the-art (SOTA) mean episode lengths of 4.52 on the CALVIN ABCD-D benchmark and 3.79 on the ABC-D benchmark using only language-labeled expert demonstrations and a comparatively lightweight architecture.Real-robot trials on a 6-DOF manipulator confirm higher success rates than comparable methods trained on identical data.These results demonstrate that our data-augmentation-enhanced two-stage training paradigm delivers efficient, scalable, and broadly applicable performance for multimodal robotic systems.

机器人操作多模态扩散模型两阶段训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。