通过视觉扰动测试发现,简单分支扩展能提升导航效果。
Seeing is Believing? Enhancing Vision-Language Navigation using Visual Perturbations
- 用深度图、扰动视图和噪声干扰视觉输入,测试模型理解力。
- 多分支架构在三个基准上达到或超越当前最优结果。
- 方法简单通用,适合各类基于拓扑的导航模型使用。
在具身环境中,基于自然语言指令的自主导航仍是视觉语言导航(VLN)智能体的挑战。尽管近年来学习多样且细粒度的视觉环境表征取得进展,但性能提升脆弱,难以确证是否真正增强了视觉定位能力,这一局限也在相关视觉语言任务中被观察到。本文初步探究先进VLN模型是否真正理解环境视觉内容,引入不同水平的视觉扰动,包括真实深度图、扰动视图和随机噪声。实验发现,即使输入含噪,简单的分支扩展反而意外提升了导航效果。受此启发,提出一种通用的多分支架构(MBA),可深入分析分支数量与视觉质量的影响。MBA将基础代理扩展为多分支结构,每个分支处理不同视觉输入。该方法简单直接,且对基于拓扑的VLN代理无特定依赖。在R2R、REVERIE、SOON三个基准上的大量实验表明,采用最优视觉组合的MBA方法达到或超越现有最先进水平。代码已公开。
原文摘要 · Abstract (English)
Autonomous navigation guided by natural language instructions in embodied environments remains a challenge for vision-language navigation (VLN) agents. Although recent advancements in learning diverse and fine-grained visual environmental representations have shown promise, the fragile performance improvements may not conclusively attribute to enhanced visual grounding,a limitation also observed in related vision-language tasks. In this work, we preliminarily investigate whether advanced VLN models genuinely comprehend the visual content of their environments by introducing varying levels of visual perturbations. These perturbations include ground-truth depth images, perturbed views and random noise. Surprisingly, we experimentally find that simple branch expansion, even with noisy visual inputs, paradoxically improves the navigational efficacy. Inspired by these insights, we further present a versatile Multi-Branch Architecture (MBA) designed to delve into the impact of both the branch quantity and visual quality. The proposed MBA extends a base agent into a multi-branch variant, where each branch processes a different visual input. This approach is embarrassingly simple yet agnostic to topology-based VLN agents. Extensive experiments on three VLN benchmarks (R2R, REVERIE, SOON) demonstrate that our method with optimal visual permutations matches or even surpasses state-of-the-art results. The source code is available at here.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。