arXiv:2509.19480cs.ROcs.LG2025-09被引 43

让机器人同时理解语言、图像和坐标,实现多模态导航

OmniVLA: An Omni-Modal Vision-Language-Action Model for Robot Navigation

  • 用随机融合多种输入方式训练模型,提升适应能力
  • 在未见过环境中表现优异,能处理缺失信息的场景
  • 适合需要灵活导航的机器人研发与教学应用

人类在导航时可灵活理解并组合语言指令、空间坐标或视觉参考等不同目标描述。而现有机器人导航策略多基于单一模态,限制了其在真实场景中的适应性。本文提出一种训练框架,支持基于视觉的机器人导航中对多模态目标的条件控制。方法采用高容量视觉-语言-动作(VLA)主干网络,通过2D位姿、第一人称图像和自然语言三种主要目标模态及其组合进行训练,结合随机模态融合策略。该设计不仅扩大可用数据集范围,还促使策略学习更丰富的几何、语义与视觉表征。所提出的OmniVLA模型在未见过环境中展现出强泛化能力,对模态缺失具有鲁棒性,并能理解新出现的自然语言指令。实验表明,OmniVLA在跨模态任务上优于专用基线模型,且具备良好可扩展性,支持快速适配新模态与新任务。我们认为OmniVLA为构建通用、灵活的导航策略提供了关键一步,也为构建多模态机器人基础模型提供了可扩展路径。视频演示已发布,模型权重与训练代码将开源。

原文摘要 · Abstract (English)

Humans can flexibly interpret and compose different goal specifications, such as language instructions, spatial coordinates, or visual references, when navigating to a destination. In contrast, most existing robotic navigation policies are trained on a single modality, limiting their adaptability to real-world scenarios where different forms of goal specification are natural and complementary. In this work, we present a training framework for robotic foundation models that enables omni-modal goal conditioning for vision-based navigation. Our approach leverages a high-capacity vision-language-action (VLA) backbone and trains with three primary goal modalities: 2D poses, egocentric images, and natural language, as well as their combinations, through a randomized modality fusion strategy. This design not only expands the pool of usable datasets but also encourages the policy to develop richer geometric, semantic, and visual representations. The resulting model, OmniVLA, achieves strong generalization to unseen environments, robustness to scarce modalities, and the ability to follow novel natural language instructions. We demonstrate that OmniVLA outperforms specialist baselines across modalities and offers a flexible foundation for fine-tuning to new modalities and tasks. We believe OmniVLA provides a step toward broadly generalizable and flexible navigation policies, and a scalable path for building omni-modal robotic foundation models. We present videos showcasing OmniVLA performance and will release its checkpoints and training code on our project page.

机器人导航多模态视觉语言动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。