兼顾显眼与细小信息,提升图文文档问答准确率
HKRAG: Holistic Knowledge Retrieval-Augmented Generation Over Visually-Rich Documents
- 用分层掩码分别捕捉醒目与细小文本信息
- 在多个基准上零样本和有监督设置均超越现有方法
- 适合需要精细理解复杂图文内容的研究者
现有针对图文丰富文档(VRD)的多模态检索增强生成(RAG)方法往往偏向于检索显著知识(如突出文字和视觉元素),而忽略关键的细小文本知识(如小字、上下文细节)。这一局限导致检索不完整,影响生成答案的准确性与全面性。为此,我们提出HKRAG,一种全新的整体式RAG框架,旨在显式捕获并融合两类知识。该框架包含两个核心组件:(1) 基于混合掩码的整体化检索器,采用显式掩码策略分别建模显著与细小知识,确保查询相关的整体信息检索;(2) 不确定性引导的代理生成器,动态评估初始回答的不确定性,并主动决定如何整合两种不同知识流以生成最优响应。在开放域图文问答基准上的大量实验表明,HKRAG在零样本和有监督设置下均持续优于现有方法,证明了整体知识检索对图文文档理解的关键作用。
原文摘要 · Abstract (English)
Existing multimodal Retrieval-Augmented Generation (RAG) methods for visually rich documents (VRD) are often biased towards retrieving salient knowledge(e.g., prominent text and visual elements), while largely neglecting the critical fine-print knowledge(e.g., small text, contextual details). This limitation leads to incomplete retrieval and compromises the generator's ability to produce accurate and comprehensive answers. To bridge this gap, we propose HKRAG, a new holistic RAG framework designed to explicitly capture and integrate both knowledge types. Our framework features two key components: (1) a Hybrid Masking-based Holistic Retriever that employs explicit masking strategies to separately model salient and fine-print knowledge, ensuring a query-relevant holistic information retrieval; and (2) an Uncertainty-guided Agentic Generator that dynamically assesses the uncertainty of initial answers and actively decides how to integrate the two distinct knowledge streams for optimal response generation. Extensive experiments on open-domain visual question answering benchmarks show that HKRAG consistently outperforms existing methods in both zero-shot and supervised settings, demonstrating the critical importance of holistic knowledge retrieval for VRD understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。