测试多模态大模型如何提升视障者视觉理解应用的可信度与满意度。
Towards Understanding the Use of MLLM-Enabled Applications for Visual Interpretation by Blind and Low Vision People
- 通过两周日记研究,收集20名视障用户使用MLLM应用的553条记录。
- 用户对视觉解释信任度均值达3.75(满分5),满意度达4.15,高风险场景也愿意信赖。
- 研究揭示了视障人群对智能视觉辅助的深层需求,指导未来系统设计。
视障人士已广泛使用基于AI的视觉理解应用以应对日常需求。尽管这些应用有所帮助,但以往研究发现用户对其频繁错误仍不满意。近期,多模态大语言模型(MLLMs)被集成到视觉理解应用中,展现出更丰富描述性解释的潜力。然而,这一进展如何改变用户使用行为尚不明确。为填补此空白,我们开展为期两周的日记研究,20名视障人士使用我们开发的MLLM启用的视觉解释应用,共收集553条记录。本文报告了6名参与者60条日记的初步分析结果。结果显示,用户认为该应用的视觉解释具有较高可信度(均值3.75/5)和满意度(均值4.15/5),甚至在医疗用药建议等高风险场景中也表现出信任。我们计划完成全样本分析,以指导未来MLLM赋能的视觉解释系统设计。
原文摘要 · Abstract (English)
Blind and Low Vision (BLV) people have adopted AI-powered visual interpretation applications to address their daily needs. While these applications have been helpful, prior work has found that users remain unsatisfied by their frequent errors. Recently, multimodal large language models (MLLMs) have been integrated into visual interpretation applications, and they show promise for more descriptive visual interpretations. However, it is still unknown how this advancement has changed people's use of these applications. To address this gap, we conducted a two-week diary study in which 20 BLV people used an MLLM-enabled visual interpretation application we developed, and we collected 553 entries. In this paper, we report a preliminary analysis of 60 diary entries from 6 participants. We found that participants considered the application's visual interpretations trustworthy (mean 3.75 out of 5) and satisfying (mean 4.15 out of 5). Moreover, participants trusted our application in high-stakes scenarios, such as receiving medical dosage advice. We discuss our plan to complete our analysis to inform the design of future MLLM-enabled visual interpretation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。