arXiv:2508.03651cs.HCcs.AI2025-08被引 4

测试ChatGPT实时视频助盲功能,发现动态场景下表现不足。

Probing the Gaps in ChatGPT Live Video Chat for Real-World Assistance for People who are Blind or Visually Impaired

  • 通过8名视障者在真实场景中使用AI视频助手,评估其辅助能力。
  • 静态场景描述准确,但动态环境中的空间与距离信息错误频发。
  • 适合研究无障碍AI交互设计,关注安全与信任问题。

大型多模态模型的进步为视障(BVI)人群通过实时视频交互理解并参与现实世界提供了新可能。然而,这类技术在支持多样化实际辅助任务中的潜力与挑战仍不明确。本文基于对八名视障参与者开展的探索性研究,让他们在陌生室内外环境中使用2024年末发布的ChatGPT Advanced Voice with Video这一前沿实时视频AI,在物品定位、视觉地标识别等任务中进行测试。结果显示,当前实时视频AI在静态视觉场景中能有效提供指导与回答,但在动态情境下无法提供必要的实时描述。尽管存在空间和距离信息的误差,参与者仍利用所获视觉信息补充其移动策略。虽然系统因高质量语音交互被感知为类人,但对用户视觉能力的误判、幻觉、泛化回复及讨好倾向导致困惑、不信任,甚至带来潜在风险。基于结果,我们讨论了辅助视频AI代理的改进方向,包括融合额外传感能力、超越对话轮次的适时干预机制,以及生态与安全考量。

原文摘要 · Abstract (English)

Recent advancements in large multimodal models have provided blind or visually impaired (BVI) individuals with new capabilities to interpret and engage with the real world through interactive systems that utilize live video feeds. However, the potential benefits and challenges of such capabilities to support diverse real-world assistive tasks remain unclear. In this paper, we present findings from an exploratory study with eight BVI participants. Participants used ChatGPT's Advanced Voice with Video, a state-of-the-art live video AI released in late 2024, in various real-world scenarios, from locating objects to recognizing visual landmarks, across unfamiliar indoor and outdoor environments. Our findings indicate that current live video AI effectively provides guidance and answers for static visual scenes but falls short in delivering essential live descriptions required in dynamic situations. Despite inaccuracies in spatial and distance information, participants leveraged the provided visual information to supplement their mobility strategies. Although the system was perceived as human-like due to high-quality voice interactions, assumptions about users' visual abilities, hallucinations, generic responses, and a tendency towards sycophancy led to confusion, distrust, and potential risks for BVI users. Based on the results, we discuss implications for assistive video AI agents, including incorporating additional sensing capabilities for real-world use, determining appropriate intervention timing beyond turn-taking interactions, and addressing ecological and safety concerns.

视障辅助实时视频AI交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。