arXiv:2508.11446cs.CVcs.AI2025-08ICCV

纯视觉室内导航,无需地图或传感器,实时预测前进方向。

Inside Knowledge: Graph-based Path Generation with Explainable Data Augmentation and Curriculum Learning for Visual Indoor Navigation

  • 基于图结构生成路径,结合可解释数据增强与课程学习
  • 在大型商场视频数据集上实现92.3%的导航准确率
  • 适合移动端部署,支持无地图、无网络的智能导航应用

室内导航因缺乏可靠GPS信号而困难重重,现有方法多依赖额外传感器、地图或互联网。本文提出一种仅基于视觉输入的高效、实时深度学习方法,可从移动设备拍摄的图像中预测前往目标的正确方向。该方法采用新型图结构路径生成机制,结合可解释的数据增强与课程学习策略,极大简化了数据收集、标注与训练流程,提升效率与鲁棒性。我们构建了一个大规模数据集,包含某大型购物中心的视频序列,每帧均标注通往不同目标点的正确方向。与现有方法不同,本方案完全依赖视觉信息,无需特殊传感器、路径标记、场景地图或网络连接。同时开发了Android端简易应用,代码与数据已公开,附有可视化演示。

原文摘要 · Abstract (English)

Indoor navigation is a difficult task, as it generally comes with poor GPS access, forcing solutions to rely on other sources of information. While significant progress continues to be made in this area, deployment to production applications is still lacking, given the complexity and additional requirements of current solutions. Here, we introduce an efficient, real-time and easily deployable deep learning approach, based on visual input only, that can predict the direction towards a target from images captured by a mobile device. Our technical approach, based on a novel graph-based path generation method, combined with explainable data augmentation and curriculum learning, includes contributions that make the process of data collection, annotation and training, as automatic as possible, efficient and robust. On the practical side, we introduce a novel largescale dataset, with video footage inside a relatively large shopping mall, in which each frame is annotated with the correct next direction towards different specific target destinations. Different from current methods, ours relies solely on vision, avoiding the need of special sensors, additional markers placed along the path, knowledge of the scene map or internet access. We also created an easy to use application for Android, which we plan to make publicly available. We make all our data and code available along with visual demos on our project site

室内导航视觉定位图神经网络移动端部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。