构建多模态文化理解基准,提升模型对文化细节的识别能力
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
- 引入检索增强机制,结合文化相关文档提升视觉理解
- 文化检索使模型在跨文化问答和图像描述任务上平均提升6%~11%
- 适用于研究多模态推理、文化感知与检索增强系统的学者
随着视觉语言模型(VLMs)日益融入日常生活,准确理解视觉文化变得愈发关键。然而,现有模型在解析文化细微差别方面表现不足。已有研究证明检索增强生成(RAG)在纯文本场景中能有效提升文化理解,但在多模态场景中的应用仍不充分。为此,我们提出RAVENEA(Retrieval-Augmented Visual culture Understanding),一个聚焦于文化理解的多模态基准,涵盖文化导向的视觉问答(cVQA)与文化驱动的图像描述(cIC)两项任务。RAVENEA通过人工标注与排序扩展了超过11,396篇独特的维基百科文档。在七种多模态检索器与十五个VLM上的广泛评估发现:(i) 文化背景标注可提升多模态检索及下游任务表现;(ii) 使用文化感知检索增强的VLM,在cVQA和cIC任务上分别平均提升6%与11%;(iii) 不同国家的文化检索增益差异显著。这些发现揭示了当前多模态检索器与VLM在文化理解上的局限,强调需在RAG系统中强化视觉文化认知。我们相信RAVENEA为推进检索增强的视觉文化理解研究提供了宝贵资源。
原文摘要 · Abstract (English)
As vision-language models (VLMs) become increasingly integrated into daily life, the need for accurate visual culture understanding is becoming critical. Yet, these models frequently fall short in interpreting cultural nuances effectively. Prior work has demonstrated the effectiveness of retrieval-augmented generation (RAG) in enhancing cultural understanding in text-only settings, while its application in multimodal scenarios remains underexplored. To bridge this gap, we introduce RAVENEA (Retrieval-Augmented Visual culturE uNdErstAnding), a new benchmark designed to advance visual culture understanding through retrieval, focusing on two tasks: culture-focused visual question answering (cVQA) and culture-informed image captioning (cIC). RAVENEA extends existing datasets by integrating over 11,396 unique Wikipedia documents curated and ranked by human annotators. Through the extensive evaluation on seven multimodal retrievers and fifteen VLMs, RAVENEA reveals some undiscovered findings: (i) In general, cultural grounding annotations can enhance multimodal retrieval and corresponding downstream tasks. (ii) VLMs, when augmented with culture-aware retrieval, generally outperform their non-augmented counterparts (by averaging +6% on cVQA and +11% on cIC). (iii) Performance of culture-aware retrieval augmented varies widely across countries. These findings highlight the limitations of current multimodal retrievers and VLMs, underscoring the need to enhance visual culture understanding within RAG systems. We believe RAVENEA offers a valuable resource for advancing research on retrieval-augmented visual culture understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。