arXiv:2603.19229cs.ROcs.AI2026-03

首个系统评估导航模型在真实干扰下的可信度基准

NavTrust: Benchmarking Trustworthiness for Embodied Navigation

  • 统一框架下对视觉、深度和指令进行真实场景干扰
  • 7个前沿模型在干扰下性能显著下降,暴露鲁棒性短板
  • 适合关注机器人导航可靠性与抗干扰能力的研究者

具身导航分为两类:视觉-语言导航(VLN),即根据自然语言指令导航;以及目标物导航(OGN),即前往指定目标物体。然而,现有研究主要在理想条件下评估模型性能,忽视了真实场景中可能发生的输入干扰。为弥补这一空白,我们提出NavTrust,一个统一基准,系统性地在现实场景中对RGB图像、深度图和指令等输入模态施加多样化干扰,并评估其对导航性能的影响。据我们所知,NavTrust是首个在统一框架下引入丰富RGB-Depth干扰和指令变化的基准。对七个先进方法的广泛评估显示,面对真实干扰时性能大幅下降,凸显出关键的鲁棒性缺陷,并为构建更可信的具身导航系统提供了路线图。此外,我们系统评估了四种不同缓解策略以增强对RGB-Depth及指令干扰的鲁棒性。基础模型包括Uni-NaVid和ETPNav,我们在真实移动机器人上部署并观察到对干扰的鲁棒性提升。项目网站:https://navtrust.github.io/。

原文摘要 · Abstract (English)

There are two major categories of embodied navigation: Vision-Language Navigation (VLN), where agents navigate by following natural language instructions; and Object-Goal Navigation (OGN), where agents navigate to a specified target object. However, existing work primarily evaluates model performance under nominal conditions, overlooking the potential corruptions that arise in real-world settings. To address this gap, we present NavTrust, a unified benchmark that systematically corrupts input modalities, including RGB, depth, and instructions, in realistic scenarios and evaluates their impact on navigation performance. To our best knowledge, NavTrust is the first benchmark that exposes embodied navigation agents to diverse RGB-Depth corruptions and instruction variations in a unified framework. Our extensive evaluation of seven state-of-the-art approaches reveals substantial performance degradation under realistic corruptions, which highlights critical robustness gaps and provides a roadmap toward more trustworthy embodied navigation systems. Furthermore, we systematically evaluate four distinct mitigation strategies to enhance robustness against RGB-Depth and instruction corruptions. Our base models include Uni-NaVid and ETPNav. We deployed them on a real mobile robot and observed improved robustness to corruptions. The project website is: https://navtrust.github.io/.

具身导航鲁棒性评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。