新基准打破视觉捷径,让多模态问答更真实可信。
Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
- 构建含关联实体的图文问答数据集,消除图像与主语的简单对应。
- 现有模型在新数据集上性能大幅下降,暴露对视觉捷径的依赖。
- 提出多图增强检索器,有效处理复杂关联实体场景。
现有多模态知识型视觉问答(MKB-VQA)基准存在‘视觉捷径’问题,因查询图像通常匹配目标文档的主要实体。我们证明模型可仅凭视觉线索取得相近效果。为此,提出基于LLM自动构建的RETINA基准,包含12万条训练数据和2千条人工校验测试数据。该基准引入次级主体(即相关实体)作为查询对象,并配以这些相关实体的图像,从而消除视觉捷径。在RETINA上评估时,现有模型性能显著下降,证实其依赖捷径。此外,提出多图多模态检索器MIMIR,通过整合多个相关实体图像丰富文档嵌入,有效应对RETINA挑战,优于仅使用单张图像的先前方法。实验验证了现有基准局限性及RETINA与MIMIR的有效性。
原文摘要 · Abstract (English)
Existing Multimodal Knowledge-Based Visual Question Answering (MKB-VQA) benchmarks suffer from "visual shortcuts", as the query image typically matches the primary subject entity of the target document. We demonstrate that models can exploit these shortcuts, achieving comparable results using visual cues alone. To address this, we introduce Relational Entity Text-Image kNowledge Augmented (RETINA) benchmark, automatically constructed using an LLM-driven pipeline, consisting of 120k training and 2k human-curated test set. RETINA contains queries referencing secondary subjects (i.e. related entities) and pairs them with images of these related entities, removing the visual shortcut. When evaluated on RETINA existing models show significantly degraded performance, confirming their reliance on the shortcut. Furthermore, we propose Multi-Image MultImodal Retriever (MIMIR), which enriches document embeddings by augmenting images of multiple related entities, effectively handling RETINA, unlike prior work that uses only a single image per document. Our experiments validate the limitations of existing benchmarks and demonstrate the effectiveness of RETINA and MIMIR. Our project is available at: Project Page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。