arXiv:2512.02631cs.LG2025-12被引 2

通过视觉提示与步骤优化,显著提升视觉语言导航成功率

SeeNav-Agent: Enhancing Vision-Language Navigation with Visual Prompt and Step-Level Policy Optimization

  • 引入双视角视觉提示减少感知幻觉,增强空间理解
  • 提出步骤级奖励分组优化法,成功率达86.7%(零样本)
  • 训练更稳定高效,适合需要精准导航的智能体应用

基于大视觉语言模型的视觉语言导航(VLN)代理常因感知、推理和规划误差导致性能受限。本文提出新框架SeeNav-Agent:首先在输入空间引入双视角视觉提示(VP),降低视觉模块的感知幻觉,提升对当前空间状态的理解;其次设计一种新的步骤级强化微调方法——步骤奖励分组策略优化(SRGPO),通过定义可验证的过程奖励并随机分组导航步骤,实现高效的步骤级优势估计。SRGPO为强化学习提供密集奖励信号,增强规划能力。在EmbodiedBench导航基准测试中,引入零样本VP模块后,GPT-4.1导航成功率达86.7%,超越当前最优LVLM约20个百分点。基于SRGPO后训练的Qwen2.5-VL-3B模型达72.3%成功率,优于最佳现有模型5.6个百分点。相比GRPO、GiGPO等算法,SRGPO在训练稳定性、收敛效率与泛化能力上均有显著提升。

原文摘要 · Abstract (English)

Existing Vision-Language Navigation (VLN) agents based on Large Vision-Language Models (LVLMs) often suffer from perception errors, reasoning errors, and planning errors, which significantly hinder their navigation performance. To address these limitations, a novel VLN agent framework, named SeeNav-Agent, is proposed in this work. First, to reduce perception hallucinations of the visual module of the VLN agent, a dual-view Visual Prompt (VP) technique is introduced in the input space, which can also improve the agent's understanding of current spatial states. Subsequently, a novel step-level Reinforcement Fine-Tuning (RFT) method, Step Reward Group Policy Optimization (SRGPO), is designed for the post-training of VLN agents. In SRGPO, we first define verifiable process rewards for the navigation task, and then perform efficient step-level advantage estimation by randomly grouping different navigation steps. SRGPO provides dense reward signals for the reinforcement learning process of the VLN agent and enhances its planning capability. Experimental results on the EmbodiedBench Navigation benchmark indicate that by introducing the zero-shot VP module, the GPT-4.1 achieves a navigation success rate of 86.7%, surpassing the current best LVLM by approximately 20 percentage points (pp). Through post-training based on SRGPO, the Qwen2.5-VL-3B model reaches a navigation success rate of 72.3%, outperforming the best existing LVLM model by 5.6 pp. Moreover, compared to RFT algorithms such as GRPO and GiGPO, the proposed SRGPO demonstrates significant improvements in training stability, convergence efficiency, and generalization capability.

视觉导航提示工程强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。