arXiv:2607.23181cs.CV2026-07

提出精简视觉语言导航框架,用最少必要信息提升智能体泛化能力。

Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation

论文配图:Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation
图 1 · 摘自论文原文
  • 基于最小充分性原则,通过选择性感知与压缩建模过滤无关信息。
  • 在R2R-CE和RxR-CE上分别提升2.0%/1.0%和0.94%/0.78%的成功率/路径长度得分。
  • 适合追求高效、鲁棒视觉语言导航的开发者与研究者。

连续环境中的视觉语言导航(VLN-CE)要求智能体将语言指令与第一人称观测对齐,并在未见场景中规划路径。尽管近期多模态大模型与世界模型方法有所进展,但常保留过多任务无关细节,削弱泛化能力并增加计算开销。本文提出BrainNav框架,遵循最小充分性原则,包含三个模块:逻辑锚定模型实现指令感知的选择性感知,抑制环境噪声;简约约束对齐模块作为紧凑的跨模态瓶颈,高效同步离散语言意图与连续潜在动态,过滤冗余信息;压缩世界模型在低秩潜在空间中预测动作条件状态。这些模块使语义意图与空间感知对齐,增强智能体在复杂任务中的鲁棒性与效率。实验表明,BrainNav在R2R-CE val-unseen上取得2.0% / 1.0%的SR/SPL提升,在RxR-CE val-unseen上提升0.94% / 0.78%,证明最小充分的世界表征是实现鲁棒视觉语言导航的有效基础。

原文摘要 · Abstract (English)

Vision-and-Language Navigation in continuous environments (VLN-CE) requires an agent to ground language in egocentric observations and plan in unseen scenes. Although recent multimodal large models and world-model-based methods have improved navigation, they often preserve excessive task-irrelevant detail, weakening generalization and increasing computational burden. We propose BrainNav, a navigation framework grounded in the Principle of Minimal Sufficiency. BrainNav consists of three components: a Logical Anchor Model that implements instruction-aware selective perception to suppress environmental noise, a Minimalist Constraint Alignment module that serves as a compact cross-modal bottleneck, efficiently synchronizing discrete linguistic intent with continuous latent dynamics while filtering out redundant information, and a Compression World Model that predicts action-conditioned states within a condensed, low-rank latent space. These modules align semantic intent with spatial perception, enhancing the agent's robustness and efficiency in complex tasks. Experiments show that BrainNav improves over prior SOTA by 2.0 % / 1.0 in SR/SPL on R2R-CE val-unseen and 0.94 % / 0.78 on RxR-CE val-unseen. These results indicate that minimally sufficient world representations provide an effective foundation for robust VLN.

视觉语言导航世界模型最小充分性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。