arXiv:2507.20892cs.RO2025-07

用拓扑图结合模型预测,让机器人仅靠视觉就能高效导航。

PixelNav: Towards Model-based Vision-Only Navigation with Topological Graphs

  • 分层架构融合视觉定位与模型预测控制
  • 真实场景实验验证系统可扩展且高效
  • 适合追求可解释性的机器人导航研究者

本文提出一种新型混合方法,实现仅依赖视觉的移动机器人导航。当前主流的端到端数据驱动模型虽灵活适应性强,但需大量训练数据且可解释性差。为此,我们构建了一种分层系统,融合模型预测控制、可通行性估计、视觉场景识别和位姿估计技术,以拓扑图表示目标环境。该方法显著提升系统可解释性,同时保持良好可扩展性。大量真实世界实验表明,所提方法在复杂环境中具备高效导航能力。

原文摘要 · Abstract (English)

This work proposes a novel hybrid approach for vision-only navigation of mobile robots, which combines advances of both deep learning approaches and classical model-based planning algorithms. Today, purely data-driven end-to-end models are dominant solutions to this problem. Despite advantages such as flexibility and adaptability, the requirement of a large amount of training data and limited interpretability are the main bottlenecks for their practical applications. To address these limitations, we propose a hierarchical system that utilizes recent advances in model predictive control, traversability estimation, visual place recognition, and pose estimation, employing topological graphs as a representation of the target environment. Using such a combination, we provide a scalable system with a higher level of interpretability compared to end-to-end approaches. Extensive real-world experiments show the efficiency of the proposed method.

视觉导航拓扑图模型预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。