发现多模态大模型常忽视视觉信息,偏信内部常识。
Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs
- 构建自动化框架+人工校验,生成视觉与常识冲突的测试数据。
- 9个模型中约20%问题过度依赖常识,尤其在是/否和动作类问题上。
- 提出'聚焦视觉'提示策略,可缓解但未根治冲突问题。
本文研究多模态大语言模型(MLLMs)在常识层面的视觉-知识冲突问题,即视觉信息与模型内部常识相矛盾的情况。为此,我们设计了一个结合人工质检的自动化框架,生成用于模拟和评估此类冲突的输入数据。基于该框架,我们构建了一个诊断基准,包含374张原创图像和1,122组高质量问答对,涵盖两类冲突和三种问题类型,具备全面评估能力。我们用该基准评估了九种代表性MLLMs的冲突解决能力,结果表明约20%的查询存在明显过度依赖参数化知识的现象,尤其集中在是/否类和动作相关问题上。在此基础上,我们评估了现有缓解方法的有效性,并对比了提出的'Focus-on-Vision'提示策略。尽管有一定改善,视觉-知识冲突仍普遍存在,且可通过我们的数据构建框架进一步放大。本研究的框架、基准与分析为理解并缓解MLLM中的视觉-知识冲突提供了重要支持。
原文摘要 · Abstract (English)
This paper explores the problem of commonsense level vision-knowledge conflict in Multimodal Large Language Models (MLLMs), where visual information contradicts model's internal commonsense knowledge. To study this issue, we introduce an automated framework, augmented with human-in-the-loop quality control, to generate inputs designed to simulate and evaluate these conflicts in MLLMs. Using this framework, we have crafted a diagnostic benchmark consisting of 374 original images and 1,122 high-quality question-answer (QA) pairs. The benchmark covers two aspects of conflict and three question types, providing a thorough assessment tool. We apply this benchmark to assess the conflict-resolution capabilities of nine representative MLLMs from various model families. Our results indicate an evident over-reliance on parametric knowledge for approximately 20% of all queries, especially among Yes-No and action-related problems. Based on these findings, we evaluate the effectiveness of existing approaches to mitigating the conflicts and compare them to our "Focus-on-Vision" prompting strategy. Despite some improvement, the vision-knowledge conflict remains unresolved and can be further scaled through our data construction framework. Our proposed framework, benchmark, and analysis contribute to the understanding and mitigation of vision-knowledge conflicts in MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。