首个专为多模态知识图谱检索设计的基准,解决大模型生成中的信息不准问题。
MKG-RAG-Bench: Benchmarking Retrieval in Multimodal Knowledge Graph-Augmented Generation

- 构建跨领域多模态知识图谱数据集,支持精准检索评估
- 实测显示检索质量直接影响生成结果准确性
- 适合研究多模态知识融合与检索的学者使用
在知识图谱增强的生成任务中,检索是关键瓶颈:多模态知识异构性强、跨模态对齐困难,且现有检索器多针对非结构化文本。为此,我们提出MKG-RAG-Bench,一个跨通用与医疗领域的多模态知识图谱基准。该基准基于两个多模态知识图谱构建,包含经严格对齐的问答数据集,支持对检索与下游生成的独立评估。通过基于大模型的清洗与标注流程,系统覆盖多种模态组合,生成结构化查询并保留精确监督信号。大量实验表明,有效多模态检索虽具挑战性,却是端到端MKG-RAG性能的关键;检索质量显著决定生成效果。本基准将检索作为首要评估目标,为诊断当前局限与推动技术进步提供坚实基础。
原文摘要 · Abstract (English)
Retrieval-augmented generation (RAG) over knowledge graphs has emerged as a promising approach for grounding large language models, yet existing benchmarks largely overlook the challenges of retrieval in multimodal knowledge graph RAG (MKG-RAG). In practice, retrieval is a critical bottleneck: multimodal knowledge is heterogeneous, difficult to align across modalities, and often poorly served by retrievers designed for unstructured corpora. To address this gap, we introduce MKG-RAG-Bench, a cross-domain benchmark explicitly designed to evaluate retrieval in MKG-RAG. MKG-RAG-Bench is constructed from two multimodal knowledge graphs spanning general and medical domains, and includes carefully aligned question-answering datasets that support controlled evaluation of both retrieval and downstream generation. The benchmark is built using an LLM-based curation pipeline that filters low-utility knowledge, generates structurally grounded queries with exact supervision, and systematically covers diverse modality configurations. Through extensive experiments across representative retriever families and modality settings, we show that effective multimodal retrieval remains challenging yet crucial for end-to-end MKG-RAG performance, and that retrieval quality strongly determines generation outcomes. By isolating retrieval as a first-class evaluation target, MKG-RAG-Bench provides a principled foundation for diagnosing current limitations and advancing multimodal knowledge graph RAG systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。