arXiv:2502.04223cs.CV2025-02被引 9

Éclair能精准提取文档文本、布局和阅读顺序,助力大模型训练与信息检索。

Éclair -- Extracting Content and Layout with Integrated Reading Order for Documents

  • 基于图像输入,联合提取文本、位置框与语义标签,保持阅读顺序。
  • 在自建多类型文档基准上达到当前最优准确率,优于现有方法。
  • 适合需要结构化文档数据的AI训练与智能文档处理场景。

光学字符识别(OCR)技术广泛用于从文档图像中提取文本,推动数字化与数据检索。然而,仅提取文本不足以理解复杂文档。全面解析需掌握其结构——包括格式、公式、表格、跨页多区块与多栏的阅读顺序,以及语义信息以识别脚注、图注等元素。这一能力对信息检索、文档问答及训练大语言模型(LLMs)和视觉语言模型(VLMs)至关重要。为此,我们提出Éclair,一种通用型文档文本提取工具,可处理多种文档类型。给定图像后,Éclair能提取按阅读顺序排列的格式化文本,同时输出对应的边界框与语义类别。为全面评估其新能力,我们构建了涵盖多样文档类型的真人标注基准。Éclair在此基准上表现优异,多项关键指标超越现有方法。此外,在多个主流基准上也展现出卓越性能,验证其通用性与强大实力。

原文摘要 · Abstract (English)

Optical Character Recognition (OCR) technology is widely used to extract text from images of documents, facilitating efficient digitization and data retrieval. However, merely extracting text is insufficient when dealing with complex documents. Fully comprehending such documents requires an understanding of their structure -- including formatting, formulas, tables, and the reading order of multiple blocks and columns across multiple pages -- as well as semantic information for detecting elements like footnotes and image captions. This comprehensive understanding is crucial for downstream tasks such as retrieval, document question answering, and data curation for training Large Language Models (LLMs) and Vision Language Models (VLMs). To address this, we introduce Éclair, a general-purpose text-extraction tool specifically designed to process a wide range of document types. Given an image, Éclair is able to extract formatted text in reading order, along with bounding boxes and their corresponding semantic classes. To thoroughly evaluate these novel capabilities, we introduce our diverse human-annotated benchmark for document-level OCR and semantic classification. Éclair achieves state-of-the-art accuracy on this benchmark, outperforming other methods across key metrics. Additionally, we evaluate Éclair on established benchmarks, demonstrating its versatility and strength across several evaluation standards.

文档理解OCR多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。