arXiv:2503.11702cs.AIcs.CL2025-03被引 6

用大模型理解室内地图,自动生成智能导航指令。

LLM-Guided Indoor Navigation with Multimodal Map Understanding

  • 利用大模型分析地图图像,生成自然语言导航指令。
  • 平均正确率达86.59%,最高达97.14%。
  • 适合智能助手、无障碍导航等场景。

室内导航因布局复杂且缺乏GNSS信号而面临挑战。现有方案常难以适应上下文,且需专用硬件。本文探索使用大型语言模型(如ChatGPT)从室内地图图像中生成自然、上下文感知的导航指令。我们在不同真实环境设计并评估了测试案例,分析大模型在解析空间布局、处理用户约束和规划高效路径方面的表现。结果表明,大模型在个性化室内导航中具有潜力,平均正确指示率达86.59%,最高达97.14%。该系统实现了高精度与强推理能力,对人工智能驱动的导航与辅助技术具有重要意义。

原文摘要 · Abstract (English)

Indoor navigation presents unique challenges due to complex layouts and the unavailability of GNSS signals. Existing solutions often struggle with contextual adaptation, and typically require dedicated hardware. In this work, we explore the potential of a Large Language Model (LLM), i.e., ChatGPT, to generate natural, context-aware navigation instructions from indoor map images. We design and evaluate test cases across different real-world environments, analyzing the effectiveness of LLMs in interpreting spatial layouts, handling user constraints, and planning efficient routes. Our findings demonstrate the potential of LLMs for supporting personalized indoor navigation, with an average of 86.59% correct indications and a maximum of 97.14%. The proposed system achieves high accuracy and reasoning performance. These results have key implications for AI-driven navigation and assistive technologies.

室内导航大模型多模态理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。