arXiv:2603.09740cs.ROcs.CV2026-03被引 2

通过分步对比对齐,让智能体在复杂环境中更准地导航

Let's Reward Step-by-Step: Step-Aware Contrastive Alignment for Vision-Language Navigation in Continuous Environments

  • 分步评估路径进展,从错误轨迹中提取有效信息
  • 在多个基准上达到当前最优性能,显著提升导航成功率
  • 适合研究视觉语言导航与强化学习融合的学者

连续环境中的视觉语言导航(VLN-CE)要求智能体从长时程人类交互中学习复杂推理。尽管多模态大语言模型(MLLM)推动了近期进展,现有训练范式难以平衡泛化能力、错误恢复与训练稳定性。具体而言,(i) 基于监督微调(SFT)的策略易产生误差累积,难以从分布外状态恢复;(ii) 强化学习微调(RFT)方法如GRPO受限于稀疏结果奖励,其二元反馈无法为单步分配信用,导致失败主导批次中梯度信号崩溃。为此,我们提出分步感知对比对齐(SACA),一种从不完美轨迹中提取密集监督的框架。核心是感知接地的分步评估器,可逐步判断进展,将失败轨迹解耦为有效前缀与精确偏差点。利用这些信号,场景条件分组构建机制动态调度批次至特定重采样与优化策略。在VLN-CE基准上的大量实验表明,SACA实现了当前最优性能。

原文摘要 · Abstract (English)

Vision-Language Navigation in Continuous Environments (VLN-CE) requires agents to learn complex reasoning from long-horizon human interactions. While Multi-modal Large Language Models (MLLMs) have driven recent progress, current training paradigms struggle to balance generalization capability, error recovery and training stability. Specifically, (i) policies derived from SFT suffer from compounding errors, struggling to recover from out-of-distribution states, and (ii) Reinforcement Fine-Tuning (RFT) methods e.g. GRPO are bottlenecked by sparse outcome rewards. Their binary feedback fails to assign credit to individual steps, leading to gradient signal collapse in failure dominant batches. To address these challenges, we introduce Step-Aware Contrastive Alignment (SACA), a framework designed to extract dense supervision from imperfect trajectories. At its core, the Perception-Grounded Step-Aware auditor evaluates progress step-by-step, disentangling failed trajectories into valid prefixes and exact divergence points. Leveraging these signals, Scenario-Conditioned Group Construction mechanism dynamically routes batches to specialized resampling and optimization strategies. Extensive experiments on VLN-CE benchmarks demonstrate that SACA achieves state-of-the-art performance.

视觉语言导航强化学习分步评估对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。