arXiv:2509.18672cs.HCcs.AI2025-09被引 3

盲人用手机找东西,NaviSense靠语音+AR+激光雷达实时导航

NaviSense: A Multimodal Assistive Mobile application for Object Retrieval by Persons with Visual Impairment

  • 通过自然语言描述物体,结合视觉-语言模型实现开放世界识别
  • 利用激光雷达与增强现实提供实时音频与触觉空间指引
  • 12名视障用户测试显示效率更高,适合日常物品查找场景

视障人士在环境中定位和取物面临巨大挑战。现有辅助技术存在权衡:高精度引导系统通常需预先扫描或仅支持固定类别物体;而具备开放世界识别能力的系统又缺乏精准的空间反馈。为此,我们提出NaviSense,一款融合对话式AI、视觉-语言模型、增强现实(AR)与LiDAR的移动辅助系统,支持通过自然语言指定目标物体,并提供无需预设的实时音觉-触觉空间引导。基于初步研究设计并经12名盲人及低视力用户验证,NaviSense显著缩短了物体检索时间,且更受用户青睐,证明了将开放世界感知与精准可访问引导相结合的价值。

原文摘要 · Abstract (English)

People with visual impairments often face significant challenges in locating and retrieving objects in their surroundings. Existing assistive technologies present a trade-off: systems that offer precise guidance typically require pre-scanning or support only fixed object categories, while those with open-world object recognition lack spatial feedback for reaching the object. To address this gap, we introduce 'NaviSense', a mobile assistive system that combines conversational AI, vision-language models, augmented reality (AR), and LiDAR to support open-world object detection with real-time audio-haptic guidance. Users specify objects via natural language and receive continuous spatial feedback to navigate toward the target without needing prior setup. Designed with insights from a formative study and evaluated with 12 blind and low-vision participants, NaviSense significantly reduced object retrieval time and was preferred over existing tools, demonstrating the value of integrating open-world perception with precise, accessible guidance.

无障碍多模态AR导航视觉-语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。