arXiv:2507.08458cs.CVcs.AI2025-07

为文档识别设计结构化归纳偏置,提升复杂文档的端到端识别能力。

A document is worth a structured record: Principled inductive bias design for document recognition

  • 将文档识别视为从文档到结构化记录的转录任务,利用内在结构分组学习。
  • 在乐谱、图形和工程图上验证,首次实现机械工程图的端到端无限制图结构识别。
  • 适用于少样本或复杂结构文档,助力未来文档基础模型统一设计。

许多文档类型依赖于内在的、约定驱动的结构来编码精确信息,如工程图的规范。然而,当前主流方法将文档识别视为纯计算机视觉问题,忽视了文档类型的结构性特征,导致依赖次优启发式后处理,使罕见或复杂文档难以被现代系统识别。本文提出新视角:将文档识别看作从文档到结构化记录的转录过程,基于转录中固有的结构对文档进行自然分组,使相关文档类型可共享学习。我们提出一种针对特定结构的关系归纳偏置设计方法,并构建适配多种结构的基线Transformer架构。在单音乐谱、形状图和简化工程图等逐步复杂的记录结构上进行大量实验,验证了该归纳偏置的有效性。通过引入无限制图结构的归纳偏置,训练出首个成功实现机械工程图端到端转录的模型,其输出为内在互连的信息。该方法对非标准文档识别(如非传统OCR、OMR)具有指导意义,有助于统一未来文档基础模型的设计。

原文摘要 · Abstract (English)

Many document types use intrinsic, convention-driven structures that serve to encode precise and structured information, such as the conventions governing engineering drawings. However, many state-of-the-art approaches treat document recognition as a mere computer vision problem, neglecting these underlying document-type-specific structural properties, making them dependent on sub-optimal heuristic post-processing and rendering many less frequent or more complicated document types inaccessible to modern document recognition. We suggest a novel perspective that frames document recognition as a transcription task from a document to a record. This implies a natural grouping of documents based on the intrinsic structure inherent in their transcription, where related document types can be treated (and learned) similarly. We propose a method to design structure-specific relational inductive biases for the underlying machine-learned end-to-end document recognition systems, and a respective base transformer architecture that we successfully adapt to different structures. We demonstrate the effectiveness of the so-found inductive biases in extensive experiments with progressively complex record structures from monophonic sheet music, shape drawings, and simplified engineering drawings. By integrating an inductive bias for unrestricted graph structures, we train the first-ever successful end-to-end model to transcribe mechanical engineering drawings to their inherently interlinked information. Our approach is relevant to inform the design of document recognition systems for document types that are less well understood than standard OCR, OMR, etc., and serves as a guide to unify the design of future document foundation models.

文档识别结构化归纳偏置工程图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。