arXiv:2602.02590cs.RO2026-02中稿 · ICRA被引 3

用结构化先验提升视觉导航的效率与安全性

StepNav: Structured Trajectory Priors for Efficient and Multimodal Visual Navigation

  • 基于几何感知概率场构建多模式路径先验
  • 生成路径更安全高效,步数减少显著
  • 适合需实时可靠导航的自动驾驶场景

视觉导航是自主系统的核心,但在复杂不确定环境中生成可靠轨迹仍是关键挑战。现有生成模型依赖无结构噪声先验,常产生不安全、低效或单模态规划,难以满足实时需求。本文提出StepNav框架,通过变分原理学习几何感知的成功概率场,识别所有可行导航通道,并构建显式的多模态混合先验,初始化条件流匹配过程。该优化被建模为带显式平滑性与安全约束的最优控制问题。通过用物理合理的候选路径替代无结构噪声,StepNav在显著更少步骤内生成更安全高效的路径。仿真与真实世界基准测试均显示其在鲁棒性、效率和安全性上优于当前最先进生成规划器,推动了实用自主导航中的可靠轨迹生成。代码已开源:https://github.com/LuoXubo/StepNav。

原文摘要 · Abstract (English)

Visual navigation is fundamental to autonomous systems, yet generating reliable trajectories in cluttered and uncertain environments remains a core challenge. Recent generative models promise end-to-end synthesis, but their reliance on unstructured noise priors often yields unsafe, inefficient, or unimodal plans that cannot meet real-time requirements. We propose StepNav, a novel framework that bridges this gap by introducing structured, multimodal trajectory priors derived from variational principles. StepNav first learns a geometry-aware success probability field to identify all feasible navigation corridors. These corridors are then used to construct an explicit, multi-modal mixture prior that initializes a conditional flow-matching process. This refinement is formulated as an optimal control problem with explicit smoothness and safety regularization. By replacing unstructured noise with physically-grounded candidates, StepNav generates safer and more efficient plans in significantly fewer steps. Experiments in both simulation and real-world benchmarks demonstrate consistent improvements in robustness, efficiency, and safety over state-of-the-art generative planners, advancing reliable trajectory generation for practical autonomous navigation. The code has been released at https://github.com/LuoXubo/StepNav.

视觉导航轨迹生成多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。