评测顶尖视觉导航模型在真实世界的表现,发现其存在碰撞频发、定位不准等系统性缺陷。
Can Vision Foundation Models Navigate? Zero-Shot Real-World Evaluation and Lessons Learned
- 在真实机器人上测试5种前沿视觉导航模型,结合路径与视觉目标识别评估
- 模型在复杂环境中碰撞率高,对相似场景区分能力弱,分布外性能显著下降
- 适用于关注机器人导航鲁棒性与真实部署的科研人员和工程师
视觉导航模型(VNMs)通过大规模视觉示范学习,有望实现通用机器人导航。尽管日益应用于真实场景,现有评估仍仅依赖成功率,掩盖了轨迹质量、碰撞行为及环境变化下的鲁棒性。我们对五种先进VNMs(GNM、ViNT、NoMaD、NaviBridger、CrossFormer)在两种机器人平台、五个室内外环境中的表现进行了真实世界评估。除成功率外,还结合路径指标与视觉目标识别得分,并通过运动模糊、阳光眩光等可控图像扰动评估鲁棒性。分析揭示三大系统性缺陷:(a) 即使是基于扩散模型和变压器架构的复杂模型也频繁发生碰撞,表明几何理解能力有限;(b) 模型无法区分感知相似但语义不同的位置,导致重复环境中的目标预测错误;(c) 在分布外条件下性能显著退化。我们将公开评估代码与数据集,以促进VNMs的可复现基准测试。
原文摘要 · Abstract (English)
Visual Navigation Models (VNMs) promise generalizable, robot navigation by learning from large-scale visual demonstrations. Despite growing real-world deployment, existing evaluations rely almost exclusively on success rate, whether the robot reaches its goal, which conceals trajectory quality, collision behavior, and robustness to environmental change. We present a real-world evaluation of five state-of-the-art VNMs (GNM, ViNT, NoMaD, NaviBridger, and CrossFormer) across two robot platforms and five environments spanning indoor and outdoor settings. Beyond success rate, we combine path-based metrics with vision-based goal-recognition scores and assess robustness through controlled image perturbations (motion blur, sunflare). Our analysis uncovers three systematic limitations: (a) even architecturally sophisticated diffusion and transformer-based models exhibit frequent collisions, indicating limited geometric understanding; (b) models fail to discriminate between different locations that are perceptually similar, however some semantics differences are present, causing goal prediction errors in repetitive environments; and (c) performance degrades under distribution shift. We will publicly release our evaluation codebase and dataset to facilitate reproducible benchmarking of VNMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。