让AI读懂含图表的多文档问答,准确率提升12%-20%。
VisDoM: Multi-Document QA with Visually Rich Elements Using Multimodal Retrieval-Augmented Generation
- 融合图文检索与生成,分步推理并约束模态一致性。
- 在跨模态信息分布场景下,准确率比基线高12-20%。
- 适合需要处理报告、幻灯片等图文混合内容的场景。
理解包含丰富视觉元素(如表格、图表、演示文稿)的多文档集合中的信息,对文档驱动的问答任务至关重要。本文提出VisDoMBench,首个全面评估多文档场景下富含多模态内容的问答系统基准。我们引入VisDoMRAG,一种新颖的多模态检索增强生成方法,同时利用视觉与文本检索,结合强大的视觉检索能力与精细的语言推理。VisDoMRAG采用多步推理流程,涵盖证据筛选与思维链推理,支持并行的文本与视觉检索生成管道。其关键创新在于推理时模态间的一致性约束融合机制,使不同模态的推理过程对齐,生成连贯答案。这提升了跨模态关键信息分布场景下的准确性,并通过隐式上下文归因增强答案可验证性。通过对开源与专有大语言模型的广泛实验,在VisDoMBench上对当前最先进的文档问答方法进行基准测试。结果表明,VisDoMRAG在端到端多模态文档问答中,相比单模态和长上下文大模型基线,准确率提升12-20%。
原文摘要 · Abstract (English)
Understanding information from a collection of multiple documents, particularly those with visually rich elements, is important for document-grounded question answering. This paper introduces VisDoMBench, the first comprehensive benchmark designed to evaluate QA systems in multi-document settings with rich multimodal content, including tables, charts, and presentation slides. We propose VisDoMRAG, a novel multimodal Retrieval Augmented Generation (RAG) approach that simultaneously utilizes visual and textual RAG, combining robust visual retrieval capabilities with sophisticated linguistic reasoning. VisDoMRAG employs a multi-step reasoning process encompassing evidence curation and chain-of-thought reasoning for concurrent textual and visual RAG pipelines. A key novelty of VisDoMRAG is its consistency-constrained modality fusion mechanism, which aligns the reasoning processes across modalities at inference time to produce a coherent final answer. This leads to enhanced accuracy in scenarios where critical information is distributed across modalities and improved answer verifiability through implicit context attribution. Through extensive experiments involving open-source and proprietary large language models, we benchmark state-of-the-art document QA methods on VisDoMBench. Extensive results show that VisDoMRAG outperforms unimodal and long-context LLM baselines for end-to-end multimodal document QA by 12-20%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。