arXiv:2608.07548cs.ROcs.CV2026-08中稿 · ICML

提出闭环自修正世界模型,提升视觉语言导航的鲁棒性

SC$^{2}$-WM: A Self-Correcting World Model with Closed-Loop Feedback for Vision-and-Language Navigation in Continuous Environments

论文配图:SC$^{2}$-WM: A Self-Correcting World Model with Closed-Loop Feedback for Vision-and-Language Navigation in Continuous Environments
图 1 · 摘自论文原文
  • 用世界模型预测反馈,动态优化导航计划
  • 测试时可选择性更新模型,应对能力不足场景
  • 在标准数据集上显著提升导航泛化能力

视觉-语言导航在连续环境(VLN-CE)中要求智能体在部分可观测条件下做出精细决策。然而,现有方法多依赖开环执行,缺乏推理过程中检测与纠正内部状态漂移的机制。本文提出SC²-WM,一种引入内部反馈的闭环决策世界模型框架。该方法通过世界模型的前瞻预测生成反馈,在执行动作前进行状态级计划修正。针对复杂场景,进一步引入条件世界感知适应机制,当反馈表明模型能力不足时,可在测试阶段选择性地更新世界模型以实现模型级修正。在标准VLN-CE基准上的实验表明,该方法显著提升了导航的鲁棒性与泛化性能。

原文摘要 · Abstract (English)

Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to make fine-grained navigation decisions under partial observability. However, most existing methods rely on open-loop execution, lacking mechanisms to detect and correct internal state drift during inference. We propose SC$^{2}$-WM, a self-correcting world model framework that introduces internal feedback for closed-loop decision making in VLN-CE. Our method derives feedback from world-model foresight to perform state-level plan refinement before action execution. To handle challenging scenarios, we further introduce conditional world-aware adaptation, which enables model-level correction by selectively updating the world model at test time when feedback indicates model capacity insufficiency. Experiments on standard VLN-CE benchmarks demonstrate improved navigation robustness and generalization. Our code is available at https://github.com/sunrise-ikun/SC2_WM.

视觉语言导航世界模型闭环控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。