arXiv:2509.25687cs.RO2025-09被引 41

统一解决导航与探索任务,支持多类型指令和实时部署。

OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation

  • 采用快慢双模块架构,分别处理短期视觉与长期规划。
  • 在多个基准上达到顶尖表现,真实场景中可稳定运行于5Hz。
  • 融合大规模通用数据提升泛化能力,适合复杂环境机器人应用。

具身导航是智能机器人面临的核心挑战,需理解视觉环境、自然语言指令并自主探索。现有模型难以在不同导航范式间提供统一解决方案,导致成功率低且泛化能力差。我们提出OmniNav,一个统一框架,涵盖指令目标、物体目标、点目标导航及基于前哨的探索。其轻量级低延迟策略能精准预测连续空间的航点(坐标与朝向),精度优于动作分块方法,支持高达5Hz的实时控制频率。架构上采用快慢系统:快速模块利用短时视野和子任务生成航点,慢速模块结合长时观测与候选前哨进行深思熟虑规划,选择后续子目标与子任务。两者协作提升路径效率,保持轨迹连贯性,尤其在探索与记忆密集场景中表现优异。关键发现是瓶颈并非仅在于导航策略学习,而在于对通用指令与物体的稳健理解。为此,OmniNav整合图像描述与视觉识别等大规模通用训练数据,形成联合多任务训练机制,显著提升成功率与鲁棒性。大量实验验证其在多种导航基准上的最先进性能,真实世界部署进一步证实其有效性。OmniNav为具身导航提供实用洞见,指明通往通用、可扩展机器人智能的可行路径。

原文摘要 · Abstract (English)

Embodied navigation presents a core challenge for intelligent robots, requiring the comprehension of visual environments, natural language instructions, and autonomous exploration. Existing models often fall short in offering a unified solution across diverse navigation paradigms, resulting in low success rates and limited generalization. We introduce OmniNav, a unified framework addressing instruct-goal, object-goal, point-goal navigation, and frontier-based exploration within a single architecture. Our approach features a lightweight, low-latency policy that accurately predicts continuous-space waypoints (coordinates and orientations). This policy surpasses action-chunk methods in precision and supports real-world deployment at control frequencies up to 5 Hz. Architecturally, OmniNav employs a fast-slow system design: a fast module generates waypoints using short-horizon visual context and subtasks, while a slow module performs deliberative planning with long-horizon observations and candidate frontiers to select subsequent subgoals and subtasks. This collaboration enhances path efficiency and maintains trajectory coherence, particularly in exploration and memory-intensive scenarios. Crucially, we identify that the primary bottleneck isn't merely navigation policy learning, but a robust understanding of general instructions and objects. To boost generalization, OmniNav integrates large-scale, general-purpose training datasets, including those for image captioning and visual recognition, into a joint multi-task regimen. This significantly improves success rates and robustness. Extensive experiments confirm OmniNav's state-of-the-art performance across various navigation benchmarks, with real-world deployment further validating its efficacy. OmniNav provides practical insights for embodied navigation, charting a scalable path towards versatile, highly generalizable robotic intelligence.

具身导航多任务学习实时控制机器人智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。