arXiv:2508.08524cs.HCcs.AI2025-08中稿 · UIST'25被引 16

让视障者通过AI语音导航体验街景,首个多模态无障碍地图工具。

StreetReaderAI: Making Street View Accessible Using Context-Aware Multimodal AI

  • 融合上下文感知与多模态AI,支持语音交互与无障碍操作
  • 覆盖2200亿张图像、100多个国家的街景数据
  • 适合视障用户远程探查地点与规划路线

交互式街景地图工具如Google Street View(GSV)和Meta Mapillary使用户可通过360°影像虚拟漫游真实环境,但对视障人群仍不可用。我们提出StreetReaderAI,首个面向视障者的可访问街景工具,结合上下文感知多模态AI、无障碍导航控制与对话式语音交互。用户可通过StreetReaderAI虚拟考察目的地、进行开放世界探索,或浏览全球超过2200亿张图像及100多个部署GSV的国家。我们与跨视觉能力团队协作迭代设计,并对11位视障用户进行了评估。结果表明,该工具在支持兴趣点(POI)调查与远程路线规划方面具有显著价值。最后,我们提出了未来研究的关键指导原则。

原文摘要 · Abstract (English)

Interactive streetscape mapping tools such as Google Street View (GSV) and Meta Mapillary enable users to virtually navigate and experience real-world environments via immersive 360° imagery but remain fundamentally inaccessible to blind users. We introduce StreetReaderAI, the first-ever accessible street view tool, which combines context-aware, multimodal AI, accessible navigation controls, and conversational speech. With StreetReaderAI, blind users can virtually examine destinations, engage in open-world exploration, or virtually tour any of the over 220 billion images and 100+ countries where GSV is deployed. We iteratively designed StreetReaderAI with a mixed-visual ability team and performed an evaluation with eleven blind users. Our findings demonstrate the value of an accessible street view in supporting POI investigations and remote route planning. We close by enumerating key guidelines for future work.

无障碍多模态AI街景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。