arXiv:2506.21556cs.CL2025-06被引 2

首个融合图文音的多模态知识图谱,助力大模型更精准推理。

VAT-KG: Knowledge-Intensive Multimodal Knowledge Graph Dataset for Retrieval-Augmented Generation

  • 构建跨模态对齐的图文音知识图谱,支持概念级知识存储。
  • 在多模态问答任务中显著提升大模型性能,准确率超基线15%以上。
  • 适合研究多模态大模型、知识增强生成与跨模态理解的学者使用。

多模态知识图谱(MMKG)通过整合多模态显式知识,弥补多模态大语言模型(MLLMs)的隐式知识不足,实现更可信的推理。然而现有MMKG普遍覆盖范围有限:多基于已有知识图谱扩展,导致知识陈旧不全,且仅支持文本与视觉等少数模态。这限制了其在视频、音频等新模态任务中的应用。为此,我们提出首个以概念为中心、知识密集型的图文音知识图谱(VAT-KG),涵盖视觉、音频与文本信息,每个三元组均关联多模态数据并附有详细概念描述。通过严格过滤与对齐流程,确保多模态数据与细粒度语义一致,可自动从任意多模态数据集生成MMKG。我们还设计新型多模态检索增强生成框架,能根据任意模态查询检索概念级知识。在多种模态的问答任务上实验表明,VAT-KG显著提升MLLM表现,验证其在统一和利用多模态知识方面的实用价值。

原文摘要 · Abstract (English)

Multimodal Knowledge Graphs (MMKGs), which represent explicit knowledge across multiple modalities, play a pivotal role by complementing the implicit knowledge of Multimodal Large Language Models (MLLMs) and enabling more grounded reasoning via Retrieval Augmented Generation (RAG). However, existing MMKGs are generally limited in scope: they are often constructed by augmenting pre-existing knowledge graphs, which restricts their knowledge, resulting in outdated or incomplete knowledge coverage, and they often support only a narrow range of modalities, such as text and visual information. These limitations restrict applicability to multimodal tasks, particularly as recent MLLMs adopt richer modalities like video and audio. Therefore, we propose the Visual-Audio-Text Knowledge Graph (VAT-KG), the first concept-centric and knowledge-intensive multimodal knowledge graph that covers visual, audio, and text information, where each triplet is linked to multimodal data and enriched with detailed descriptions of concepts. Specifically, our construction pipeline ensures cross-modal knowledge alignment between multimodal data and fine-grained semantics through a series of stringent filtering and alignment steps, enabling the automatic generation of MMKGs from any multimodal dataset. We further introduce a novel multimodal RAG framework that retrieves detailed concept-level knowledge in response to queries from arbitrary modalities. Experiments on question answering tasks across various modalities demonstrate the effectiveness of VAT-KG in supporting MLLMs, highlighting its practical value in unifying and leveraging multimodal knowledge.

多模态知识图谱检索增强生成大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。