评测多模态大模型做视障人士视觉助手的效果与不足
Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users
- 基于用户调研设计五项真实任务,含新型光学盲文识别
- 12个模型均在文化语境、多语言支持上表现不佳
- 适合关注无障碍AI的开发者与研究者参考
本文探讨多模态大语言模型(MLLMs)作为视障人士辅助技术的有效性。通过用户调查,发现尽管采用率高,但用户仍面临上下文理解、文化敏感性及复杂场景解析等挑战,尤其对依赖模型进行视觉解读的群体影响显著。基于此,我们设计了五项以图像和视频输入为主的用户中心任务,包括一项新的光学盲文识别任务。对12个MLLMs的系统评估显示,其在文化背景理解、多语言支持、盲文读写、辅助物体识别及幻觉问题上仍需改进。本研究为多模态AI在无障碍领域的未来发展提供关键洞见,强调构建更具包容性、鲁棒性和可信度的视觉辅助技术的必要性。
原文摘要 · Abstract (English)
This paper explores the effectiveness of Multimodal Large Language models (MLLMs) as assistive technologies for visually impaired individuals. We conduct a user survey to identify adoption patterns and key challenges users face with such technologies. Despite a high adoption rate of these models, our findings highlight concerns related to contextual understanding, cultural sensitivity, and complex scene understanding, particularly for individuals who may rely solely on them for visual interpretation. Informed by these results, we collate five user-centred tasks with image and video inputs, including a novel task on Optical Braille Recognition. Our systematic evaluation of twelve MLLMs reveals that further advancements are necessary to overcome limitations related to cultural context, multilingual support, Braille reading comprehension, assistive object recognition, and hallucinations. This work provides critical insights into the future direction of multimodal AI for accessibility, underscoring the need for more inclusive, robust, and trustworthy visual assistance technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。