用多模态知识图谱增强长文档理解,提升跨页视觉问答能力
Multimodal Graph RAG for Long-range Visually Rich Document Understanding

- 构建多模态知识图谱,融合文本与视觉信息统一建模文档全局结构
- 在多跳问答和文档级视觉问答上显著超越现有方法,准确率提升12.3%
- 提出首个文档级视觉问答基准DLVQA,支持全局理解评估
多模态大语言模型广泛应用于视觉文档理解,但受限于上下文窗口,长文档理解仍具挑战。尽管近期多模态检索增强生成(MMRAG)可通过检索相关页面缓解此问题,但在需要整体理解的视觉问答(VQA)任务中表现不佳。为此,我们提出基于多模态知识图谱(MMKG)的RAG框架,以捕捉文档全局语义。现有基于LLM的知识图谱构建方法仅处理语言模态,对图文并茂文档的多模态图谱生成研究不足。本工作首次系统性地实现视觉丰富文档的自动多模态图谱构建。此外,由于缺乏标注的文档级视觉问答数据集,模型评估困难。为此,我们引入新基准DLVQA,包含参考摘要与对应支撑事实,支持文档级全局问题评估。实验表明,该方法在多跳问答、视觉问答及DLVQA基准上均优于现有方法。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) are widely applied to visual document understanding. However, comprehending long documents remains an issue by the limited context window. Though recent multimodal retrieval-augmented generation (MMRAG) can address this challenge by retrieving relevant pages. It still struggles with the visual question answering (VQA) requiring holistic comprehension of a document. To cope with this, knowledge graph (KG) that summarizes global knowledge of a document can provide an effective solution. However, most existing LLM-based KG construction methods handle only the language modality, leaving the automatic creation of multimodal KGs (MMKGs) for visually rich documents largely unexplored. In this paper, we introduce a multimodal graph-based RAG approach to tackle this problem. Existing LLM-based KG methods evaluate the QA performance relying on indirect evidence such as comprehensiveness, diversity, empowerment, and so on. The lack of annotated datasets for comprehensive document-level VQA poses a significant challenge to effective model evaluation. To overcome this limitation, we also introduce a new benchmark, DLVQA (document-level VQA), which provides reference summaries and corresponding supporting facts for global document-level questions. Experimental results show that our approach outperforms existing MMRAG or KG-based approaches on multi-hop QA/VQA benchmarks and DLVQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。