arXiv:2511.01210cs.CVcs.RO2025-11中稿 · ICRA被引 10

让机器人用多传感器看懂物理世界,提升抓取成功率。

OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation

  • 用多模态感知融合红外、雷达和麦克风信息,生成统一空间标记图像。
  • 真实场景下平均任务成功率84%,比纯视觉模型高59%。
  • 适合需要环境感知的机器人操作,如工业装配或家庭服务。

视觉-语言-动作(VLA)模型通过大规模视觉-语言预训练展现了强大的动作预测泛化能力。然而,现有模型大多仅依赖RGB摄像头,限制了感知范围与操作能力。我们提出OmniVLA,一种整合新型传感模态的全模态VLA模型,实现超越RGB的物理空间智能。核心是传感器掩码图像:将红外、毫米波雷达和麦克风阵列提供的空间定位与物理意义掩码叠加于RGB图像上,形成统一表征。该设计保持输入与RGB统计一致,便于训练,提供跨硬件的统一接口,并支持轻量级每传感器投影器实现高效学习。基于此,构建多感官视觉-语言-动作模型架构,并在已预训练的RGB-VLA主干基础上进行训练。在复杂真实任务中评估显示,传感器模态感知引导机器人操作,OmniVLA平均任务成功率达84%,显著优于纯RGB模型(+59%)和原始传感器输入基线(+28%),同时具备更高学习效率与更强泛化能力。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models have shown strong generalization for robotic action prediction through large-scale vision-language pretraining. However, most existing models rely solely on RGB cameras, limiting their perception and, consequently, manipulation capabilities. We present OmniVLA, an omni-modality VLA model that integrates novel sensing modalities for physically-grounded spatial intelligence beyond RGB perception. The core of our approach is the sensor-masked image, a unified representation that overlays spatially grounded and physically meaningful masks onto the RGB images, derived from sensors including an infrared camera, a mmWave radar, and a microphone array. This image-native unification keeps sensor input close to RGB statistics to facilitate training, provides a uniform interface across sensor hardware, and enables data-efficient learning with lightweight per-sensor projectors. Built on this, we present a multisensory vision-language-action model architecture and train the model based on an RGB-pretrained VLA backbone. We evaluate OmniVLA on challenging real-world tasks where sensor-modality perception guides the robotic manipulation. OmniVLA achieves an average task success rate of 84%, significantly outperforms both RGB-only and raw-sensor-input baseline models by 59% and 28% respectively, meanwhile showing higher learning efficiency and stronger generalization capability.

机器人操作多模态感知物理空间智能视觉-语言-动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。