arXiv:2508.08120cs.LGcs.AI2025-08

用手机摄像头和大模型实现无需基站的室内定位与导航

Vision-Based Localization and LLM-based Navigation for Indoor Environments

  • 用微调的ResNet-50分析手机拍摄图像定位位置
  • 定位准确率达96%,导航指令正确率75%
  • 适合医院、机场等缺乏基础设施的场所

室内导航因缺乏可靠GPS信号及复杂建筑结构而面临挑战。本文提出一种结合视觉定位与大语言模型(LLM)导航的方法。定位系统采用两阶段微调的ResNet-50卷积神经网络,通过智能手机摄像头输入识别用户位置。导航模块则利用大语言模型,基于精心设计的系统提示,解析预处理后的楼层平面图并生成分步指引。在具有重复特征且可视范围受限的真实办公室走廊中进行实验,模型在所有测试路径点均达到96%的高置信度定位准确率,即使在视野受限和短时查询条件下依然稳定。使用ChatGPT对真实建筑平面图进行导航测试,平均指令准确率为75%,但存在零样本推理能力不足和推理延迟问题。研究证明,仅依赖现成摄像头和公开楼层图即可实现可扩展的无基础设施室内导航,尤其适用于医院、机场、教育机构等资源受限场景。

原文摘要 · Abstract (English)

Indoor navigation remains a complex challenge due to the absence of reliable GPS signals and the architectural intricacies of large enclosed environments. This study presents an indoor localization and navigation approach that integrates vision-based localization with large language model (LLM)-based navigation. The localization system utilizes a ResNet-50 convolutional neural network fine-tuned through a two-stage process to identify the user's position using smartphone camera input. To complement localization, the navigation module employs an LLM, guided by a carefully crafted system prompt, to interpret preprocessed floor plan images and generate step-by-step directions. Experimental evaluation was conducted in a realistic office corridor with repetitive features and limited visibility to test localization robustness. The model achieved high confidence and an accuracy of 96% across all tested waypoints, even under constrained viewing conditions and short-duration queries. Navigation tests using ChatGPT on real building floor maps yielded an average instruction accuracy of 75%, with observed limitations in zero-shot reasoning and inference time. This research demonstrates the potential for scalable, infrastructure-free indoor navigation using off-the-shelf cameras and publicly available floor plans, particularly in resource-constrained settings like hospitals, airports, and educational institutions.

室内导航视觉定位大模型应用手机智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。