用手机拍图实现无障碍室内导航,无需安装设备或人工标记。
NaVIP: An Image-Centric Indoor Navigation Solution for Visually Impaired People
- 基于手机拍摄的图像构建无基础设施导航系统。
- 采集30万张带6自由度位姿标签的图像,支持精准定位与环境理解。
- 适合无障碍设计、智能助盲研究者,可直接用于开发实用导航工具。
室内导航因缺乏卫星定位而困难重重,对视障人士(VIPs)尤为严峻,因其无法获取导视信息。现有基于蓝牙、激光雷达等传感器的导航方案需部署大量标签或昂贵硬件,且依赖大量人工投入,难以规模化。本文提出一种无基础设施、任务可扩展的图像中心化导航解决方案NaVIP,以促进视觉智能在无障碍场景中的应用。我们首先在四层科研楼中采集大规模手机摄像头数据,共30万张图像,每张图像均标注精确的6自由度相机位姿、室内兴趣点(PoIs)信息及描述性标题,以辅助视障用户理解环境。在两个核心方面进行基准测试:1)定位系统性能;2)探索支持能力,重点提升训练可扩展性与实时推理效率。所构建的数据集、代码与模型权重已公开于https://github.com/junfish/VIP_Navi。
原文摘要 · Abstract (English)
Indoor navigation is challenging due to the absence of satellite positioning. This challenge is manifold greater for Visually Impaired People (VIPs) who lack the ability to get information from wayfinding signage. Other sensor signals (e.g., Bluetooth and LiDAR) can be used to create turn-by-turn navigation solutions with position updates for users. Unfortunately, these solutions require tags to be installed all around the environment or the use of fairly expensive hardware. Moreover, these solutions require a high degree of manual involvement that raises costs, thus hampering scalability. We propose an image dataset and associated image-centric solution called NaVIP towards visual intelligence that is infrastructure-free and task-scalable, and can assist VIPs in understanding their surroundings. Specifically, we start by curating large-scale phone camera data in a four-floor research building, with 300K images, to lay the foundation for creating an image-centric indoor navigation and exploration solution for inclusiveness. Every image is labelled with precise 6DoF camera poses, details of indoor PoIs, and descriptive captions to assist VIPs. We benchmark on two main aspects: 1) positioning system and 2) exploration support, prioritizing training scalability and real-time inference, to validate the prospect of image-based solution towards indoor navigation. The dataset, code, and model checkpoints are made publicly available at https://github.com/junfish/VIP_Navi.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。