让导航机器人共享视觉观察,提升任务表现。
Does Peer Observation Help? Vision-Sharing Collaboration for Vision-Language Navigation
- 多机器人在相同环境间交换看到的视觉记忆,扩展感知范围。
- 在R2R基准上,两种模型均显著提升,最优达+18.3%成功率。
- 适合研究协作式智能体导航的学者与开发者。
视觉-语言导航(VLN)系统受限于局部观测,只能基于自身访问过的位置积累知识。随着多个机器人共存于同一环境,一个自然问题浮现:协同导航的智能体能否从彼此的观测中获益?本文提出Co-VLN,一种极简、模型无关的框架,用于系统探究并发导航的智能体共享观测是否有益。当独立导航的智能体识别出共同经过的位置时,它们会交换结构化的感知记忆,从而在不增加探索成本的情况下扩展各自的感受野。我们在R2R基准上,针对学习型DUET和零样本MapGPT两种代表性范式验证了该框架,并开展大量分析实验,系统揭示了同伴观测共享在VLN中的内在机制。结果表明,启用视觉共享的模型在两种范式下均取得显著性能提升,为未来协作具身导航研究奠定了坚实基础。
原文摘要 · Abstract (English)
Vision-Language Navigation (VLN) systems are fundamentally constrained by partial observability, as an agent can only accumulate knowledge from locations it has personally visited. As multiple robots increasingly coexist in shared environments, a natural question arises: can agents navigating the same space benefit from each other's observations? In this work, we introduce Co-VLN, a minimalist, model-agnostic framework for systematically investigating whether and how peer observations from concurrently navigating agents can benefit VLN. When independently navigating agents identify common traversed locations, they exchange structured perceptual memory, effectively expanding each agent's receptive field at no additional exploration cost. We validate our framework on the R2R benchmark under two representative paradigms (the learning-based DUET and the zero-shot MapGPT), and conduct extensive analytical experiments to systematically reveal the underlying dynamics of peer observation sharing in VLN. Results demonstrate that vision-sharing enabled model yields substantial performance improvements across both paradigms, establishing a strong foundation for future research in collaborative embodied navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。