arXiv:2602.13469cs.HCcs.AI2026-02被引 2

盲人用户用大模型看图,发现问答虽准但常出错或沉默。

How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision People

  • 让用户通过对话提问获取视觉信息,突破传统描述局限。
  • 22.2%回答错误,10.8%拒绝回应,可靠性待提升。
  • 提出‘视觉助手’技能,指导应用更好服务视障人群。

多模态大语言模型(MLLMs)正在改变视障人士获取视觉信息的方式。与仅提供描述的传统工具不同,基于MLLM的应用支持对话式交互,用户可提问以获取目标相关细节。然而,其在真实场景中的表现及对视障者日常生活的影响仍缺乏证据。为此,我们开展为期两周的日记研究,记录20名视障参与者使用一款MLLM赋能的视觉解释应用的情况。尽管用户对应用生成的视觉解释评价为‘可信’(均值3.76/5)和‘较满意’(均值4.13/5),但人工智能仍存在22.2%的错误回答或10.8%的回避回应。研究显示,尽管MLLM能提升描述准确性,但真正支持日常使用还需具备‘视觉助手’能力——即以目标为导向、可靠的服务行为。论文据此提出‘视觉助手’技能及应用设计指南,助力提升视障人士对视觉信息的可及性。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) are changing how Blind and Low Vision (BLV) people access visual information. Unlike traditional visual interpretation tools that only provide descriptions, MLLM-enabled applications offer conversational assistance, where users can ask questions to obtain goal-relevant details. However, evidence about their performance in the real-world and implications for BLV people's daily lives remains limited. To address this, we conducted a two-week diary study, where we captured 20 BLV participants' use of an MLLM-enabled visual interpretation application. Although participants rated the visual interpretations of the application as "trustworthy" (mean=3.76 out of 5, max=extremely trustworthy) and "somewhat satisfying" (mean=4.13 out of 5, max=very satisfying), the AI often produced incorrect answers (22.2%) or abstained (10.8%) from responding to users' requests. Our findings show that while MLLMs can improve visual interpretations' descriptive accuracy, supporting everyday use also depends on the "visual assistant" skill: behaviors for providing goal-directed, reliable assistance. We conclude by proposing the "visual assistant" skill and guidelines to help MLLM-enabled visual interpretation applications better support BLV people's access to visual information.

多模态视障辅助大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。