arXiv:2508.07989cs.CVcs.HC2025-08ICCV被引 5

AI助盲系统常看不清自动扶梯方向,存在隐性运动盲区

The Escalator Problem: Identifying Implicit Motion Blindness in AI for Accessibility

  • 指出当前多模态模型在视频理解中因帧采样导致运动感知缺陷
  • 实验证明顶尖模型无法识别自动扶梯运行方向
  • 呼吁构建更安全、贴近真实需求的辅助技术评估标准

多模态大语言模型在辅助视障人士方面潜力巨大,但我们发现其在实际应用中存在关键缺陷。本文提出‘自动扶梯问题’——即最先进模型无法感知自动扶梯的运行方向,作为‘隐性运动盲区’这一深层局限性的典型例证。该盲区源于视频理解中以离散静态图像序列为主的主流采样范式,难以捕捉连续且信号微弱的运动。本文并非提出新模型,而是:(I) 明确阐述这一盲点;(II) 分析其对用户信任的影响;(III) 发出变革呼吁。我们主张从单纯语义识别转向稳健的物理感知,并倡导开发以安全、可靠及用户真实需求为核心的新型人本化基准。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) hold immense promise as assistive technologies for the blind and visually impaired (BVI) community. However, we identify a critical failure mode that undermines their trustworthiness in real-world applications. We introduce the Escalator Problem -- the inability of state-of-the-art models to perceive an escalator's direction of travel -- as a canonical example of a deeper limitation we term Implicit Motion Blindness. This blindness stems from the dominant frame-sampling paradigm in video understanding, which, by treating videos as discrete sequences of static images, fundamentally struggles to perceive continuous, low-signal motion. As a position paper, our contribution is not a new model but rather to: (I) formally articulate this blind spot, (II) analyze its implications for user trust, and (III) issue a call to action. We advocate for a paradigm shift from purely semantic recognition towards robust physical perception and urge the development of new, human-centered benchmarks that prioritize safety, reliability, and the genuine needs of users in dynamic environments.

AI助盲运动感知多模态可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。