arXiv:2506.21316cs.CV2025-06被引 1

多粒度文档视觉定位,精准回答复杂多语言文档问题

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents

  • 融合多语言OCR与大模型,实现块/行/词/点四级定位
  • 线级粒度下准确率最优,比现有方法提升显著
  • 适用于需要高精度解释的多语言文档问答场景

文本密集型文档图像中的视觉定位是文档智能与视觉问答系统的关键挑战。本文提出DRISHTIKON,一种支持多粒度、多块的视觉定位框架,旨在提升复杂多语言文档中VQA系统的可解释性与可信度。该方法结合多语言OCR、大语言模型与新型区域匹配算法,在块、行、词、点四个层级实现答案区域定位。我们构建了多粒度视觉定位(MGVG)基准数据集,涵盖多个行业发布的多样化圆形通知,每条数据均经人工精细标注并验证,覆盖多种粒度标签。大量实验表明,本方法在定位准确率上达到当前最佳水平,其中线级粒度在精确率与召回率间取得最佳平衡。消融实验进一步证实多块与多行推理的优势。对比评估显示,主流视觉语言模型在精确定位方面表现不佳,凸显了本方法结构化对齐策略的有效性。研究成果为真实世界文本主导场景下的鲁棒、可解释文档理解提供了新路径。代码与数据集已公开。

原文摘要 · Abstract (English)

Visual grounding in text-rich document images is a critical yet underexplored challenge for Document Intelligence and Visual Question Answering (VQA) systems. We present DRISHTIKON, a multi-granular and multi-block visual grounding framework designed to enhance interpretability and trust in VQA for complex, multilingual documents. Our approach integrates multilingual OCR, large language models, and a novel region matching algorithm to localize answer spans at the block, line, word, and point levels. We introduce the Multi-Granular Visual Grounding (MGVG) benchmark, a curated test set of diverse circular notifications from various sectors, each manually annotated with fine-grained, human-verified labels across multiple granularities. Extensive experiments show that our method achieves state-of-the-art grounding accuracy, with line-level granularity providing the best balance between precision and recall. Ablation studies further highlight the benefits of multi-block and multi-line reasoning. Comparative evaluations reveal that leading vision-language models struggle with precise localization, underscoring the effectiveness of our structured, alignment-based approach. Our findings pave the way for more robust and interpretable document understanding systems in real-world, text-centric scenarios with multi-granular grounding support. Code and dataset are made available for future research.

视觉定位文档理解多粒度多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。