arXiv:2409.15250cs.CVcs.RO2024-09ICRA被引 29

解决机器人模型视觉泛化不足问题,提升跨场景表现

ReVLA: Reverting Visual Domain Limitation of Robotic Foundation Models

  • 提出渐进式骨干反转方法,缓解视觉灾难性遗忘
  • 在视觉域外任务中抓取与提升性能分别提升77%和66%
  • 适合关注机器人通用性与视觉鲁棒性的研究者

近期大型语言模型和大规模机器人数据集的发展推动了机器人模型向通用化转变,涌现出一批开放的视觉-语言-动作模型,在多种任务中表现出色。本文研究了三种现有机器人基础模型的视觉泛化能力,发现其在视觉域外(OOD)场景下缺乏鲁棒性,可能源于训练数据变化有限或灾难性遗忘。我们进一步分析了使用两个预训练视觉基础模型的OpenVLA,发现其在深度回归任务中因DINO-v2出现灾难性遗忘而失败。为此,我们提出基于模型融合的渐进式骨干反转方法,使OpenVLA重新获得视觉泛化能力。ReVLA在视觉域外抓取和提升任务上分别相较OpenVLA提升77%和66%。完整评估、回放轨迹及模型权重已公开于ReVLA页面。

原文摘要 · Abstract (English)

Recent progress in large language models and access to large-scale robotic datasets has sparked a paradigm shift in robotics models transforming them into generalists able to adapt to various tasks, scenes, and robot modalities. A large step for the community are open Vision Language Action models which showcase strong performance in a wide variety of tasks. In this work, we study the visual generalization capabilities of three existing robotic foundation models, and propose a corresponding evaluation framework. Our study shows that the existing models do not exhibit robustness to visual out-of-domain scenarios. This is potentially caused by limited variations in the training data and/or catastrophic forgetting, leading to domain limitations in the vision foundation models. We further explore OpenVLA, which uses two pre-trained vision foundation models and is, therefore, expected to generalize to out-of-domain experiments. However, we showcase catastrophic forgetting by DINO-v2 in OpenVLA through its failure to fulfill the task of depth regression. To overcome the aforementioned issue of visual catastrophic forgetting, we propose a gradual backbone reversal approach founded on model merging. This enables OpenVLA -- which requires the adaptation of the visual backbones during initial training -- to regain its visual generalization ability. Regaining this capability enables our ReVLA model to improve over OpenVLA by a factor of 77\% and 66\% for grasping and lifting in visual OOD tasks. Comprehensive evaluations, episode rollouts and model weights are available on the ReVLA Page

机器人视觉泛化灾难性遗忘模型融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。