arXiv:2507.15692cs.HCcs.CL2025-07被引 11

通过展示多个模型描述的差异,帮助视障用户发现错误图像描述。

Surfacing Variations to Calibrate Perceived Reliability of MLLM-generated Image Descriptions

  • 用多种方式呈现多模型输出差异,辅助判断可靠性
  • 用户识别不可靠描述的能力提升4.9倍
  • 15位参与者中14人更喜欢看差异信息

多模态大语言模型(MLLM)为盲人和低视力(BLV)人群提供了获取视觉信息的新途径。然而,这些模型常产生难以察觉的错误,可能在药品识别、穿搭选择等场景中带来安全与社交风险。尽管BLV用户会通过交叉核对工具或求助他人来应对,但此类方法耗时且不实用。本文研究如何系统性地呈现多个MLLM响应之间的差异,以帮助视障用户在不依赖视觉的情况下识别不可靠信息。我们提出一种设计空间,实现三种差异展示风格的原型系统,并开展包含15名BLV参与者的用户研究。结果表明,展示差异可使用户识别不可靠陈述的能力提升4.9倍,显著降低对MLLM输出的信任度。14名参与者偏好查看多模型描述差异,全部表示希望将该系统用于理解龙卷风路径、发布社交媒体图片等任务。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) provide new opportunities for blind and low vision (BLV) people to access visual information in their daily lives. However, these models often produce errors that are difficult to detect without sight, posing safety and social risks in scenarios from medication identification to outfit selection. While BLV MLLM users use creative workarounds such as cross-checking between tools and consulting sighted individuals, these approaches are often time-consuming and impractical. We explore how systematically surfacing variations across multiple MLLM responses can support BLV users to detect unreliable information without visually inspecting the image. We contribute a design space for eliciting and presenting variations in MLLM descriptions, a prototype system implementing three variation presentation styles, and findings from a user study with 15 BLV participants. Our results demonstrate that presenting variations significantly increases users' ability to identify unreliable claims (by 4.9x using our approach compared to single descriptions) and significantly decreases perceived reliability of MLLM responses. 14 of 15 participants preferred seeing variations of MLLM responses over a single description, and all expressed interest in using our system for tasks from understanding a tornado's path to posting an image on social media.

多模态模型视障辅助可靠性评估信息可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。