arXiv:2507.05805cs.CV2025-07被引 2

DREAM端到端重建文档,兼顾布局与内容,效果超越现有方法。

DREAM: Document Reconstruction via End-to-end Autoregressive Model

  • 设计端到端自回归模型,统一处理文档元素序列生成。
  • 在新构建的DocRec1K数据集上达到最优性能,优于多阶段方法。
  • 适合需要高精度文档结构还原的研究与工业应用。

文档重建是文档分析与识别领域的重要任务,近年来受到广泛关注。现有方法多采用分阶段策略,先分别预测文本、表格、公式等子任务,再通过启发式规则整合,但存在误差传播问题。尽管已有研究尝试用生成模型端到端提取文本、表格和公式逻辑顺序,却难以保留关键的版面布局信息。为此,本文提出面向文档重建的端到端自回归模型DREAM,将文档图像转化为包含丰富元素信息的统一序列。同时,我们定义了标准化的文档重建任务,构建了新的文档相似性度量(DSM)和DocRec1K数据集以评估性能。实验表明,DREAM在文档重建任务中表现卓越;在文档版面分析、文本识别、表格结构识别、公式识别及阅读顺序检测等多项子任务上也具备竞争力,展现出良好的通用性。

原文摘要 · Abstract (English)

Document reconstruction constitutes a significant facet of document analysis and recognition, a field that has been progressively accruing interest within the scholarly community. A multitude of these researchers employ an array of document understanding models to generate predictions on distinct subtasks, subsequently integrating their results into a holistic document reconstruction format via heuristic principles. Nevertheless, these multi-stage methodologies are hindered by the phenomenon of error propagation, resulting in suboptimal performance. Furthermore, contemporary studies utilize generative models to extract the logical sequence of plain text, tables and mathematical expressions in an end-to-end process. However, this approach is deficient in preserving the information related to element layouts, which are vital for document reconstruction. To surmount these aforementioned limitations, we in this paper present an innovative autoregressive model specifically designed for document reconstruction, referred to as Document Reconstruction via End-to-end Autoregressive Model (DREAM). DREAM transmutes the text image into a sequence of document reconstruction in a comprehensive, end-to-end process, encapsulating a broader spectrum of document element information. In addition, we establish a standardized definition of the document reconstruction task, and introduce a novel Document Similarity Metric (DSM) and DocRec1K dataset for assessing the performance of the task. Empirical results substantiate that our methodology attains unparalleled performance in the realm of document reconstruction. Furthermore, the results on a variety of subtasks, encompassing document layout analysis, text recognition, table structure recognition, formula recognition and reading order detection, indicate that our model is competitive and compatible with various tasks.

文档重建自回归模型版面理解端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。