智能手杖融合社交导航与多模态交互,提升视障者室内移动体验。
PISHYAR: A Socially Intelligent Smart Cane for Indoor Social Navigation and Multimodal Human-Robot Interaction for Visually Impaired People
- 基于树莓派5与OAK-D Lite,融合视觉感知与动态路径规划实现避障定位。
- 多模态交互系统在真实场景中达成约80%准确率,支持语音与视觉双模式切换。
- 用户测试显示高接受度,尤其在交互自然性与社会信任感方面表现突出。
本文提出PISHYAR,一款由研究团队设计的社交智能手杖,结合社交感知导航与多模态人机交互,同时支持物理移动与互动辅助。系统由两部分构成:(1) 基于Raspberry Pi 5的社交导航框架,集成OAK-D Lite相机的实时RGB-D感知、基于YOLOv8的物体检测、基于COMPOSER的群体活动识别、D* Lite动态路径规划,以及通过振动马达提供触觉反馈,用于寻找空座位等任务;(2) 支持语音与视觉双模式切换的代理式多模态大模型交互框架,整合语音识别、视觉语言模型(VLM)、大语言模型(LLM)与文本转语音技术,实现自然语音对话、场景描述与目标定位。系统通过仿真测试、实地实验与以用户为中心的研究进行评估。模拟与真实室内环境结果显示,在不同社交条件下系统整体准确率达约80%,障碍物规避与社交合规导航表现可靠;群体活动识别在多种人群场景下均保持稳健性能。此外,针对八名视障及低视力用户的初步探索性用户研究,通过结构化任务与UTAUT问卷,揭示了系统在可用性、信任度与社会亲和力方面的积极感知,具有较高接受度。结果表明,PISHYAR作为多模态辅助移动设备,不仅超越传统导航功能,还能提供社交互动支持。
原文摘要 · Abstract (English)
This paper presents PISHYAR, a socially intelligent smart cane designed by our group to combine socially aware navigation with multimodal human-AI interaction to support both physical mobility and interactive assistance. The system consists of two components: (1) a social navigation framework implemented on a Raspberry Pi 5 that integrates real-time RGB-D perception using an OAK-D Lite camera, YOLOv8-based object detection, COMPOSER-based collective activity recognition, D* Lite dynamic path planning, and haptic feedback via vibration motors for tasks such as locating a vacant seat; and (2) an agentic multimodal LLM-VLM interaction framework that integrates speech recognition, vision language models, large language models, and text-to-speech, with dynamic routing between voice-only and vision-only modes to enable natural voice-based communication, scene description, and object localization from visual input. The system is evaluated through a combination of simulation-based tests, real-world field experiments, and user-centered studies. Results from simulated and real indoor environments demonstrate reliable obstacle avoidance and socially compliant navigation, achieving an overall system accuracy of approximately 80% under different social conditions. Group activity recognition further shows robust performance across diverse crowd scenarios. In addition, a preliminary exploratory user study with eight visually impaired and low-vision participants evaluates the agentic interaction framework through structured tasks and a UTAUT-based questionnaire reveals high acceptance and positive perceptions of usability, trust, and perceived sociability during our experiments. The results highlight the potential of PISHYAR as a multimodal assistive mobility aid that extends beyond navigation to provide socially interactive support for such users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。