研究发现视觉语言模型在读漫画时常产生幻觉,影响残障用户理解。
Semantic Similarity is a Spurious Measure of Comic Understanding: Lessons Learned from Hallucinations in a Benchmarking Experiment
- 构建漫画理解基准,评估模型页面级解读能力
- 发现模型常虚构不存在的物体,形成系统性幻觉
- 为残障用户设计更可靠的漫画辅助系统提供方向
能帮助视障用户获取漫画/漫画内容的系统尚属空白。生成式视觉-语言模型(VLMs)在图像描述和漫画理解方面展现潜力,但现有研究多局限于单面板分析。为真正支持视障群体,需加强页面级理解与解释能力。本文提出初步的VLM漫画解读性能基准,识别并分类过程中出现的幻觉,建立通用物体幻觉分类体系。最后给出未来研究建议,强调幻觉抑制与漫画数据集质量提升的重要性。
原文摘要 · Abstract (English)
A system that enables blind or visually impaired users to access comics/manga would introduce a new medium of storytelling to this community. However, no such system currently exists. Generative vision-language models (VLMs) have shown promise in describing images and understanding comics, but most research on comic understanding is limited to panel-level analysis. To fully support blind and visually impaired users, greater attention must be paid to page-level understanding and interpretation. In this work, we present a preliminary benchmark of VLM performance on comic interpretation tasks. We identify and categorize hallucinations that emerge during this process, organizing them into generalized object-hallucination taxonomies. We conclude with guidance on future research, emphasizing hallucination mitigation and improved data curation for comic interpretation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。