arXiv:2602.06356cs.ROcs.SY2026-02被引 2

解决视觉语言导航中因偏离导致的累积误差问题。

Nipping the Drift in the Bud: Retrospective Rectification for Robust Vision-Language Navigation

  • 通过反事实回溯和决策条件监督合成,实现在线修正
  • 在R2R-CE和RxR-CE上达成最优成功率与SPL
  • 适合研究具身智能与多模态导航的学者

视觉语言导航(VLN)要求智能体根据自然语言指令在复杂连续3D环境中导航。然而,主流的模仿学习范式存在暴露偏差:推理时的微小偏离会导致误差不断累积。尽管DAgger类方法试图通过纠正错误状态缓解该问题,我们发现一个关键缺陷——指令-状态错位。强制智能体从偏离路径的状态学习恢复动作,常产生与原指令语义冲突的监督信号。为此,我们提出BudVLN,一种基于在线策略采样的框架,通过构建与当前状态分布一致的监督信号进行学习。BudVLN利用测地线查询器,通过反事实回溯和决策条件监督合成,生成源自有效历史状态的纠正轨迹,确保语义一致性。在标准R2R-CE和RxR-CE基准上的实验表明,BudVLN持续缓解分布偏移,同时在成功率达与SPL指标上达到当前最优表现。

原文摘要 · Abstract (English)

Vision-Language Navigation (VLN) requires embodied agents to interpret natural language instructions and navigate through complex continuous 3D environments. However, the dominant imitation learning paradigm suffers from exposure bias, where minor deviations during inference lead to compounding errors. While DAgger-style approaches attempt to mitigate this by correcting error states, we identify a critical limitation: Instruction-State Misalignment. Forcing an agent to learn recovery actions from off-track states often creates supervision signals that semantically conflict with the original instruction. In response to these challenges, we introduce BudVLN, an online framework that learns from on-policy rollouts by constructing supervision to match the current state distribution. BudVLN performs retrospective rectification via counterfactual re-anchoring and decision-conditioned supervision synthesis, using a geodesic oracle to synthesize corrective trajectories that originate from valid historical states, ensuring semantic consistency. Experiments on the standard R2R-CE and RxR-CE benchmarks demonstrate that BudVLN consistently mitigates distribution shift and achieves state-of-the-art performance in both Success Rate and SPL.

视觉语言导航具身智能强化学习路径修正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。