arXiv:2410.21169cs.MMcs.AI2024-10被引 67

系统梳理文档解析技术,揭示结构化信息提取的关键方法与未来方向。

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction

  • 按模块化流水线与视觉语言模型两类组织现有解析方法
  • 涵盖版面分析、多模态内容识别及评估基准体系
  • 适合研究文档智能、知识抽取与RAG系统的读者

文档解析(DP)将非结构化或半结构化文档转换为可机器读取的结构化表示,支持知识库构建和检索增强生成(RAG)等下游应用。本文综述了文档解析领域的最新进展,提出一个系统性分类框架,将现有方法分为基于模块化流水线的系统与由视觉-语言模型(VLMs)驱动的统一模型。详细回顾了流水线系统中的关键组件,包括版面分析,以及对文本、表格、数学表达式和视觉元素等异构内容的识别,并系统追踪了用于文档解析的专用VLMs的发展历程。此外,总结了广泛采用的评估指标与高质量基准数据集,确立了当前解析质量的标准。最后,讨论了若干关键开放挑战,如复杂版面的鲁棒性、VLM-based解析的可靠性及推理效率,并展望了构建更准确、可扩展的文档智能系统的发展方向。

原文摘要 · Abstract (English)

Document parsing (DP) transforms unstructured or semi-structured documents into structured, machine-readable representations, enabling downstream applications such as knowledge base construction and retrieval-augmented generation (RAG). This survey provides a comprehensive and timely review of document parsing research. We propose a systematic taxonomy that organizes existing approaches into modular pipeline-based systems and unified models driven by Vision-Language Models (VLMs). We provide a detailed review of key components in pipeline systems, including layout analysis and the recognition of heterogeneous content such as text, tables, mathematical expressions, and visual elements, and then systematically track the evolution of specialized VLMs for document parsing. Additionally, we summarize widely adopted evaluation metrics and high-quality benchmarks that establish current standards for parsing quality. Finally, we discuss key open challenges, including robustness to complex layouts, reliability of VLM-based parsing, and inference efficiency, and outline directions for building more accurate and scalable document intelligence systems.

文档解析视觉语言模型信息抽取RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。