用大模型将平面图转为可导航知识图谱,提升视障者室内导航准确率
Floorplan2Guide: LLM-Guided Floorplan Parsing for BLV Indoor Navigation
- 用大语言模型解析建筑平面图,自动提取空间信息
- 5次提示下最高达92.3%导航准确率,长路径仍保持61.5%
- 知识图谱+上下文学习让视障用户导航更精准
室内导航对视障人群仍是重大挑战。现有方案多依赖基础设施,难以适应动态环境。本文提出一种新方法,利用基础模型将平面图转化为可导航知识图谱,并生成自然语言导航指令。Floorplan2Guide通过大语言模型(LLM)从建筑布局中提取空间信息,显著减少早期解析方法所需的手动预处理。实验表明,在模拟与真实场景中,少样本学习相比零样本学习提升了导航准确率。在MP-1平面图上,采用5次提示的Claude 3.7 Sonnet达到92.31%、76.92%和61.54%的短、中、长路径准确率。基于图结构的空间推理成功率比直接视觉推理高出15.4%,验证了图形表示与上下文学习能有效提升导航性能,使本方案更适合盲人及低视力(BLV)用户。
原文摘要 · Abstract (English)
Indoor navigation remains a critical challenge for people with visual impairments. The current solutions mainly rely on infrastructure-based systems, which limit their ability to navigate safely in dynamic environments. We propose a novel navigation approach that utilizes a foundation model to transform floor plans into navigable knowledge graphs and generate human-readable navigation instructions. Floorplan2Guide integrates a large language model (LLM) to extract spatial information from architectural layouts, reducing the manual preprocessing required by earlier floorplan parsing methods. Experimental results indicate that few-shot learning improves navigation accuracy in comparison to zero-shot learning on simulated and real-world evaluations. Claude 3.7 Sonnet achieves the highest accuracy among the evaluated models, with 92.31%, 76.92%, and 61.54% on the short, medium, and long routes, respectively, under 5-shot prompting of the MP-1 floor plan. The success rate of graph-based spatial structure is 15.4% higher than that of direct visual reasoning among all models, which confirms that graphical representation and in-context learning enhance navigation performance and make our solution more precise for indoor navigation of Blind and Low Vision (BLV) users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。