arXiv:2603.15370cs.CV2026-03被引 4

用强化学习提升导航模型对指令扰动的鲁棒性。

Trajectory-Diversity-Driven Robust Vision-and-Language Navigation

  • 基于组内相对策略优化,探索多样路径提升泛化能力。
  • 在未知环境中实现R2R+3.0%、REVERIE+1.71%的SPL提升。
  • 特别适合需要高鲁棒性的真实场景导航应用。

视觉-语言导航(VLN)要求智能体根据自然语言指令在逼真环境中导航。现有方法主要依赖模仿学习,泛化能力有限且对执行扰动敏感。本文提出NavGRPO,一种基于群组相对策略优化的强化学习框架,通过探索多样化轨迹并进行组内性能比较来优化策略,使智能体能在不依赖额外价值网络的情况下区分有效策略与专家路径。在ScaleVLN基础上,NavGRPO在R2R和REVERIE基准上实现未见环境下的+3.0%和+1.71% SPL提升。在极端早期扰动下,相较基线提升+14.89% SPL,验证了目标导向强化学习能显著增强导航策略的鲁棒性。代码与模型将公开。

原文摘要 · Abstract (English)

Vision-and-Language Navigation (VLN) requires agents to navigate photo-realistic environments following natural language instructions. Current methods predominantly rely on imitation learning, which suffers from limited generalization and poor robustness to execution perturbations. We present NavGRPO, a reinforcement learning framework that learns goal-directed navigation policies through Group Relative Policy Optimization. By exploring diverse trajectories and optimizing via within-group performance comparisons, our method enables agents to distinguish effective strategies beyond expert paths without requiring additional value networks. Built on ScaleVLN, NavGRPO achieves superior robustness on R2R and REVERIE benchmarks with +3.0% and +1.71% SPL improvements in unseen environments. Under extreme early-stage perturbations, we demonstrate +14.89% SPL gain over the baseline, confirming that goal-directed RL training builds substantially more robust navigation policies. Code and models will be released.

视觉语言导航强化学习鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。