arXiv:2607.14643cs.RO2026-07中稿 · presentation at IR…

用少步生成+批评者引导优化,提升视觉导航速度与避障能力

NavCMPO: Critic-Guided MeanFlow Policy Optimization for Adaptive Navigation

论文配图:NavCMPO: Critic-Guided MeanFlow Policy Optimization for Adaptive Navigation
图 1 · 摘自论文原文
  • 先用少步生成轨迹,再由批评者梯度修正中间路径
  • 在InternVLA-N1上成功率74.7%,比基线高6.4个百分点,推理延迟降至60毫秒
  • 适合需要快速适应真实机器人的自主导航场景

端到端扩散策略在无地图视觉导航中表现优异,但其迭代去噪过程带来显著推理延迟,而行为克隆性能受限于专家示范质量。我们提出NavCMPO,一种两阶段自适应导航框架,结合少步MeanFlow轨迹生成、批评者引导精修和强化学习微调。预训练阶段,通过障碍物接近度预测任务,使视觉表征捕捉障碍感知空间信息。为补偿少步生成导致的避障性能下降,采用基于障碍点云监督训练的批评者,利用梯度对中间轨迹进行精修。适配阶段,使用带行为克隆正则化的近端策略优化对MeanFlow策略进行微调,同时更新批评者以适应本体观测变化。在InternVLA-N1基准上,相同训练预算下,NavCMPO平均成功率达74.7%,优于重训练的NavDP基线6.4个百分点,推理延迟从85毫秒降至60毫秒。在Unitree Go2上的实验进一步验证了良好的仿真到现实迁移能力。

原文摘要 · Abstract (English)

End-to-end diffusion-based policies have demonstrated strong performance in mapless visual navigation, but their iterative denoising process introduces substantial inference latency, while behavior cloning limits performance to the quality of expert demonstrations. We present NavCMPO, a two-stage adaptive navigation framework that combines few-step MeanFlow trajectory generation, critic-guided refinement, and reinforcement learning fine-tuning. During pre-training, an obstacle proximity prediction task encourages the visual representation to capture obstacle-aware spatial information. To compensate for the degradation in obstacle avoidance caused by few-step generation, Critic-Guided Trajectory Refinement (CGTR) uses gradients from a critic trained with obstacle-point-cloud supervision to refine intermediate trajectories. During adaptation, the MeanFlow policy is fine-tuned using Proximal Policy Optimization with behavior-cloning regularization, while the critic is updated to accommodate embodiment-specific observation changes. Under a matched training budget on the InternVLA-N1 benchmark, NavCMPO achieves an average success rate of 74.7\%, exceeding the retrained NavDP baseline by 6.4 percentage points, while reducing inference latency from 85\,ms to 60\,ms. Experiments on a Unitree Go2 further demonstrate effective sim-to-real transfer.

视觉导航扩散模型强化学习机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。