AI助盲系统常看不清自动扶梯方向,存在隐性运动盲区
The Escalator Problem: Identifying Implicit Motion Blindness in AI for Accessibility
- 指出当前多模态模型在视频理解中因帧采样导致运动感知缺陷
- 实验证明顶尖模型无法识别自动扶梯运行方向
- 呼吁构建更安全、贴近真实需求的辅助技术评估标准
多模态大语言模型在辅助视障人士方面潜力巨大,但我们发现其在实际应用中存在关键缺陷。本文提出‘自动扶梯问题’——即最先进模型无法感知自动扶梯的运行方向,作为‘隐性运动盲区’这一深层局限性的典型例证。该盲区源于视频理解中以离散静态图像序列为主的主流采样范式,难以捕捉连续且信号微弱的运动。本文并非提出新模型,而是:(I) 明确阐述这一盲点;(II) 分析其对用户信任的影响;(III) 发出变革呼吁。我们主张从单纯语义识别转向稳健的物理感知,并倡导开发以安全、可靠及用户真实需求为核心的新型人本化基准。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) hold immense promise as assistive technologies for the blind and visually impaired (BVI) community. However, we identify a critical failure mode that undermines their trustworthiness in real-world applications. We introduce the Escalator Problem -- the inability of state-of-the-art models to perceive an escalator's direction of travel -- as a canonical example of a deeper limitation we term Implicit Motion Blindness. This blindness stems from the dominant frame-sampling paradigm in video understanding, which, by treating videos as discrete sequences of static images, fundamentally struggles to perceive continuous, low-signal motion. As a position paper, our contribution is not a new model but rather to: (I) formally articulate this blind spot, (II) analyze its implications for user trust, and (III) issue a call to action. We advocate for a paradigm shift from purely semantic recognition towards robust physical perception and urge the development of new, human-centered benchmarks that prioritize safety, reliability, and the genuine needs of users in dynamic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。