arXiv:2511.20227cs.IR2025-11被引 7

兼顾显眼与细小信息,提升图文文档问答准确率

HKRAG: Holistic Knowledge Retrieval-Augmented Generation Over Visually-Rich Documents

  • 用分层掩码分别捕捉醒目与细小文本信息
  • 在多个基准上零样本和有监督设置均超越现有方法
  • 适合需要精细理解复杂图文内容的研究者

现有针对图文丰富文档(VRD)的多模态检索增强生成(RAG)方法往往偏向于检索显著知识(如突出文字和视觉元素),而忽略关键的细小文本知识(如小字、上下文细节)。这一局限导致检索不完整,影响生成答案的准确性与全面性。为此,我们提出HKRAG,一种全新的整体式RAG框架,旨在显式捕获并融合两类知识。该框架包含两个核心组件:(1) 基于混合掩码的整体化检索器,采用显式掩码策略分别建模显著与细小知识,确保查询相关的整体信息检索;(2) 不确定性引导的代理生成器,动态评估初始回答的不确定性,并主动决定如何整合两种不同知识流以生成最优响应。在开放域图文问答基准上的大量实验表明,HKRAG在零样本和有监督设置下均持续优于现有方法,证明了整体知识检索对图文文档理解的关键作用。

原文摘要 · Abstract (English)

Existing multimodal Retrieval-Augmented Generation (RAG) methods for visually rich documents (VRD) are often biased towards retrieving salient knowledge(e.g., prominent text and visual elements), while largely neglecting the critical fine-print knowledge(e.g., small text, contextual details). This limitation leads to incomplete retrieval and compromises the generator's ability to produce accurate and comprehensive answers. To bridge this gap, we propose HKRAG, a new holistic RAG framework designed to explicitly capture and integrate both knowledge types. Our framework features two key components: (1) a Hybrid Masking-based Holistic Retriever that employs explicit masking strategies to separately model salient and fine-print knowledge, ensuring a query-relevant holistic information retrieval; and (2) an Uncertainty-guided Agentic Generator that dynamically assesses the uncertainty of initial answers and actively decides how to integrate the two distinct knowledge streams for optimal response generation. Extensive experiments on open-domain visual question answering benchmarks show that HKRAG consistently outperforms existing methods in both zero-shot and supervised settings, demonstrating the critical importance of holistic knowledge retrieval for VRD understanding.

图文理解检索增强细粒度信息多模态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。