解耦视觉语言导航中的感知、推理与纠错,提升长程导航稳定性
DecoVLN: Decoupling Observation, Reasoning, and Correction for Vision-and-Language Navigation
- 将长期记忆构建为优化问题,动态筛选语义相关、视觉多样且时间覆盖广的帧
- 通过地学距离量化偏差,仅在可信区域收集高质量数据进行纠错训练
- 适用于复杂3D环境下的长程导航,适合需高鲁棒性的智能体系统
视觉-语言导航(VLN)要求智能体遵循长程指令并穿越复杂3D环境。现有方法面临两大挑战:构建有效的长期记忆库,以及克服误差累积问题。为此,我们提出DecoVLN框架,实现长程导航中稳健的流式感知与闭环控制。首先,将长期记忆构建建模为优化问题,引入自适应精炼机制,通过迭代优化统一评分函数从历史候选帧池中选择帧。该函数联合平衡三个关键标准:与指令的语义相关性、已选记忆间的视觉多样性,以及历史轨迹的时间覆盖度。其次,为缓解误差累积,提出状态-动作对级别的纠正微调策略。利用状态间测地距离精确量化偏离专家轨迹的程度,使智能体在可信区域内收集高质量状态-动作对,同时过滤低相关性的污染数据。这提升了纠错的效率与稳定性。大量实验验证了DecoVLN的有效性,我们已在真实环境中部署该系统。
原文摘要 · Abstract (English)
Vision-and-Language Navigation (VLN) requires agents to follow long-horizon instructions and navigate complex 3D environments. However, existing approaches face two major challenges: constructing an effective long-term memory bank and overcoming the compounding errors problem. To address these issues, we propose DecoVLN, an effective framework designed for robust streaming perception and closed-loop control in long-horizon navigation. First, we formulate long-term memory construction as an optimization problem and introduce adaptive refinement mechanism that selects frames from a historical candidate pool by iteratively optimizing a unified scoring function. This function jointly balances three key criteria: semantic relevance to the instruction, visual diversity from the selected memory, and temporal coverage of the historical trajectory. Second, to alleviate compounding errors, we introduce a state-action pair-level corrective finetuning strategy. By leveraging geodesic distance between states to precisely quantify deviation from the expert trajectory, the agent collects high-quality state-action pairs in the trusted region while filtering out the polluted data with low relevance. This improves both the efficiency and stability of error correction. Extensive experiments demonstrate the effectiveness of DecoVLN, and we have deployed it in real-world environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。