提出精简视觉语言导航框架,用最少必要信息提升智能体泛化能力。
Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation

- 基于最小充分性原则,通过选择性感知与压缩建模过滤无关信息。
- 在R2R-CE和RxR-CE上分别提升2.0%/1.0%和0.94%/0.78%的成功率/路径长度得分。
- 适合追求高效、鲁棒视觉语言导航的开发者与研究者。
连续环境中的视觉语言导航(VLN-CE)要求智能体将语言指令与第一人称观测对齐,并在未见场景中规划路径。尽管近期多模态大模型与世界模型方法有所进展,但常保留过多任务无关细节,削弱泛化能力并增加计算开销。本文提出BrainNav框架,遵循最小充分性原则,包含三个模块:逻辑锚定模型实现指令感知的选择性感知,抑制环境噪声;简约约束对齐模块作为紧凑的跨模态瓶颈,高效同步离散语言意图与连续潜在动态,过滤冗余信息;压缩世界模型在低秩潜在空间中预测动作条件状态。这些模块使语义意图与空间感知对齐,增强智能体在复杂任务中的鲁棒性与效率。实验表明,BrainNav在R2R-CE val-unseen上取得2.0% / 1.0%的SR/SPL提升,在RxR-CE val-unseen上提升0.94% / 0.78%,证明最小充分的世界表征是实现鲁棒视觉语言导航的有效基础。
原文摘要 · Abstract (English)
Vision-and-Language Navigation in continuous environments (VLN-CE) requires an agent to ground language in egocentric observations and plan in unseen scenes. Although recent multimodal large models and world-model-based methods have improved navigation, they often preserve excessive task-irrelevant detail, weakening generalization and increasing computational burden. We propose BrainNav, a navigation framework grounded in the Principle of Minimal Sufficiency. BrainNav consists of three components: a Logical Anchor Model that implements instruction-aware selective perception to suppress environmental noise, a Minimalist Constraint Alignment module that serves as a compact cross-modal bottleneck, efficiently synchronizing discrete linguistic intent with continuous latent dynamics while filtering out redundant information, and a Compression World Model that predicts action-conditioned states within a condensed, low-rank latent space. These modules align semantic intent with spatial perception, enhancing the agent's robustness and efficiency in complex tasks. Experiments show that BrainNav improves over prior SOTA by 2.0 % / 1.0 in SR/SPL on R2R-CE val-unseen and 0.94 % / 0.78 on RxR-CE val-unseen. These results indicate that minimally sufficient world representations provide an effective foundation for robust VLN.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。