arXiv:2512.10956cs.CV2025-12

用双目视觉和中间层感知提升城市动态导航效率

Empowering Dynamic Urban Navigation with Stereo and Mid-Level Vision

  • 融合双目输入与深度估计等中层视觉模块,避免纯单目端到端训练
  • 仅用1.5%数据达到当前最优水平,全量数据下性能更优
  • 适合做复杂城市环境导航的机器人系统研究者

基础模型在语言与视觉领域的成功推动了端到端机器人导航基础模型(NFMs)的研究。这些模型直接将单目视觉输入映射为控制动作,完全忽略中层视觉模块(如追踪、深度估计等)。尽管隐式学习视觉能力的假设具有吸引力,但其依赖大量像素到动作的标注数据,难以获取。在动态非结构化环境中,鲁棒导航需要精确的几何与动态理解,而单目视觉的深度尺度模糊性进一步限制了空间推理。本文表明,依赖单目视觉并忽略中层视觉先验是低效的。我们提出StereoWalker,通过引入双目输入和显式的中层视觉(如深度估计、密集像素追踪),缓解深度模糊问题,并利用现代中层视觉模型提供动态场景中的可靠几何与运动结构。我们还构建了一个大型双目导航数据集,基于互联网双目视频自动标注动作,以支持StereoWalker训练并促进后续研究。实验表明,中层视觉使StereoWalker仅用1.5%训练数据即达到当前最佳性能,并在使用全部数据时超越现有方法。此外,双目输入相比单目输入显著提升导航表现。

原文摘要 · Abstract (English)

The success of foundation models in language and vision motivated research in fully end-to-end robot navigation foundation models (NFMs). NFMs directly map monocular visual input to control actions and ignore mid-level vision modules (tracking, depth estimation, etc) entirely. While the assumption that vision capabilities will emerge implicitly is compelling, it requires large amounts of pixel-to-action supervision that are difficult to obtain. The challenge is especially pronounced in dynamic and unstructured settings, where robust navigation requires precise geometric and dynamic understanding, while the depth-scale ambiguity in monocular views further limits accurate spatial reasoning. In this paper, we show that relying on monocular vision and ignoring mid-level vision priors is inefficient. We present StereoWalker, which augments NFMs with stereo inputs and explicit mid-level vision such as depth estimation and dense pixel tracking. Our intuition is straightforward: stereo inputs resolve the depth-scale ambiguity, and modern mid-level vision models provide reliable geometric and motion structure in dynamic scenes. We also curate a large stereo navigation dataset with automatic action annotation from Internet stereo videos to support training of StereoWalker and to facilitate future research. Through our experiments, we find that mid-level vision enables StereoWalker to achieve a comparable performance as the state-of-the-art using only 1.5% of the training data, and surpasses the state-of-the-art using the full data. We also observe that stereo vision yields higher navigation performance than monocular input.

导航双目视觉中层感知机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。