arXiv:2505.14381cs.AI2025-05Conference of the …被引 9

SCAN提升图文文档检索生成效果,精准分割内容区域。

SCAN: Semantic Document Layout Analysis for Textual and Visual Retrieval-Augmented Generation

  • 用语义粒度划分文档区域,兼顾上下文与效率
  • 在英日语数据集上文本和视觉RAG分别提升9.4和10.4点
  • 适合需要高效处理复杂图文文档的生成系统

随着大语言模型(LLMs)和视觉-语言模型(VLMs)的广泛应用,用于检索增强生成(RAG)及视觉RAG的丰富文档分析技术日益受到关注。研究表明,使用VLM可提升RAG性能,但处理富含信息的文档仍具挑战,因单页包含大量内容。本文提出SCAN(SemantiC Document Layout ANalysis),一种新型方法,增强面向视觉丰富文档的文本与视觉RAG系统。该方法兼容VLM,能以适当语义粒度识别文档组件,在保留上下文与处理效率间取得平衡。SCAN采用粗粒度语义策略,将文档划分为涵盖连续组件的连贯区域。我们在标注数据集上微调目标检测模型训练SCAN。实验结果表明,在英文与日文数据集上,应用SCAN使端到端文本RAG性能最高提升9.4点,视觉RAG性能最高提升10.4点,优于传统方法甚至商用文档处理方案。

原文摘要 · Abstract (English)

With the increasing adoption of Large Language Models (LLMs) and Vision-Language Models (VLMs), rich document analysis technologies for applications like Retrieval-Augmented Generation (RAG) and visual RAG are gaining significant attention. Recent research indicates that using VLMs yields better RAG performance, but processing rich documents remains a challenge since a single page contains large amounts of information. In this paper, we present SCAN (SemantiC Document Layout ANalysis), a novel approach that enhances both textual and visual Retrieval-Augmented Generation (RAG) systems that work with visually rich documents. It is a VLM-friendly approach that identifies document components with appropriate semantic granularity, balancing context preservation with processing efficiency. SCAN uses a coarse-grained semantic approach that divides documents into coherent regions covering contiguous components. We trained the SCAN model by fine-tuning object detection models on an annotated dataset. Our experimental results across English and Japanese datasets demonstrate that applying SCAN improves end-to-end textual RAG performance by up to 9.4 points and visual RAG performance by up to 10.4 points, outperforming conventional approaches and even commercial document processing solutions.

文档分析视觉RAG语义分割多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。