用大模型+文本地图,帮视障者在室内自动导航
Guide-LLM: An Embodied LLM Agent and Text-Based Topological Map for Robotic Guidance of People with Visual Impairments
- 用文本拓扑图代替复杂地图,让大模型规划直行与转角路径
- 能识别危险并根据用户偏好定制路线,提升导航安全性与个性化
- 适合研究无障碍辅助系统或人机交互的开发者与学者
视障人士的导航面临重大挑战。传统辅助工具如盲杖和导盲犬虽有价值,但难以提供详细空间信息与精准引导。本文提出 Guide-LLM,一种基于大语言模型(LLM)的具身智能体,用于帮助视障者在大型室内环境中导航。该方法引入一种新型文本式拓扑地图,使 LLM 能以简化的环境表示进行全局路径规划,重点聚焦于直行与直角转弯,从而简化导航逻辑。同时,利用 LLM 的常识推理能力实现障碍物检测,并根据用户偏好进行个性化路径规划。模拟实验表明该系统在引导视障者方面具有显著成效,展现出高效、自适应且个性化的导航辅助潜力,为辅助技术领域带来重要进展。
原文摘要 · Abstract (English)
Navigation presents a significant challenge for persons with visual impairments (PVI). While traditional aids such as white canes and guide dogs are invaluable, they fall short in delivering detailed spatial information and precise guidance to desired locations. Recent developments in large language models (LLMs) and vision-language models (VLMs) offer new avenues for enhancing assistive navigation. In this paper, we introduce Guide-LLM, an embodied LLM-based agent designed to assist PVI in navigating large indoor environments. Our approach features a novel text-based topological map that enables the LLM to plan global paths using a simplified environmental representation, focusing on straight paths and right-angle turns to facilitate navigation. Additionally, we utilize the LLM's commonsense reasoning for hazard detection and personalized path planning based on user preferences. Simulated experiments demonstrate the system's efficacy in guiding PVI, underscoring its potential as a significant advancement in assistive technology. The results highlight Guide-LLM's ability to offer efficient, adaptive, and personalized navigation assistance, pointing to promising advancements in this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。