研究视觉语言模型中图文信息如何影响问答表现,发现互补信息提升准确率,冲突信息则导致错误。
Why context matters in VQA and Reasoning: Semantic interventions for VLM input modalities
- 通过语义干预工具操控图文输入,分析模态间关系对模型决策的影响。
- 图文互补时准确率与推理质量提升,矛盾信息使模型更不自信且出错率上升。
- 揭示图像主导模型判断,特别发现PaliGemma存在危险过自信问题。
生成式AI的幻觉与失败问题凸显了理解视觉语言模型(VLM)中多模态信息作用的重要性。本文研究图像与文本模态在视觉问答(VQA)与推理任务中的融合机制,通过答案准确率、推理质量、模型不确定性和模态相关性进行评估。我们构建了用于基准测试的语义干预-VQA(SI-VQA)数据集,并开发交互式语义干预(ISI)工具,实现对图文输入的精细操控与分析。结果表明,模态间互补信息能提升答案准确性与推理质量,而矛盾信息会损害性能并降低模型信心。图像文本标注对准确率和不确定性影响微弱,但略微提高图像相关性。注意力分析证实图像输入在VQA中占据主导地位。本研究评估了多种前沿VLM模型,发现PaliGemma存在有害的过度自信现象,比LLaVA模型更易发生无声失效。该工作为模态融合的严谨分析提供了数据与方法基础。
原文摘要 · Abstract (English)
The various limitations of Generative AI, such as hallucinations and model failures, have made it crucial to understand the role of different modalities in Visual Language Model (VLM) predictions. Our work investigates how the integration of information from image and text modalities influences the performance and behavior of VLMs in visual question answering (VQA) and reasoning tasks. We measure this effect through answer accuracy, reasoning quality, model uncertainty, and modality relevance. We study the interplay between text and image modalities in different configurations where visual content is essential for solving the VQA task. Our contributions include (1) the Semantic Interventions (SI)-VQA dataset, (2) a benchmark study of various VLM architectures under different modality configurations, and (3) the Interactive Semantic Interventions (ISI) tool. The SI-VQA dataset serves as the foundation for the benchmark, while the ISI tool provides an interface to test and apply semantic interventions in image and text inputs, enabling more fine-grained analysis. Our results show that complementary information between modalities improves answer and reasoning quality, while contradictory information harms model performance and confidence. Image text annotations have minimal impact on accuracy and uncertainty, slightly increasing image relevance. Attention analysis confirms the dominant role of image inputs over text in VQA tasks. In this study, we evaluate state-of-the-art VLMs that allow us to extract attention coefficients for each modality. A key finding is PaliGemma's harmful overconfidence, which poses a higher risk of silent failures compared to the LLaVA models. This work sets the foundation for rigorous analysis of modality integration, supported by datasets specifically designed for this purpose.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。