arXiv:2412.12902cs.CV2024-12被引 1

通过图文对齐提升文档布局分析,无需OCR也能精准理解文档图像。

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment

  • 设计图文对齐机制,让视觉模型更好利用文档中的文本信息。
  • 在D4LA和FUNSD上刷新性能纪录,且比大模型更高效。
  • 不依赖OCR,适合实际部署中无预处理文本的场景。

多模态学习的兴起显著提升了文档智能水平,文档被视为同时包含文本与视觉信息的多模态实体。然而,现有研究多聚焦于文本,将视觉信息作为辅助。部分纯视觉方法需推理时输入OCR识别的文本,或缺乏文本对齐机制。为此,本文提出专为文档图像设计的图文对齐技术——DoPTA。该模型通过此技术在无需OCR的前提下,在多项文档图像理解任务中表现优异。结合辅助重建目标,DoPTA以更少的预训练算力超越更大规模模型,在D4LA与FUNSD两个挑战性基准上取得新SOTA结果。

原文摘要 · Abstract (English)

The advent of multimodal learning has brought a significant improvement in document AI. Documents are now treated as multimodal entities, incorporating both textual and visual information for downstream analysis. However, works in this space are often focused on the textual aspect, using the visual space as auxiliary information. While some works have explored pure vision based techniques for document image understanding, they require OCR identified text as input during inference, or do not align with text in their learning procedure. Therefore, we present a novel image-text alignment technique specially designed for leveraging the textual information in document images to improve performance on visual tasks. Our document encoder model DoPTA - trained with this technique demonstrates strong performance on a wide range of document image understanding tasks, without requiring OCR during inference. Combined with an auxiliary reconstruction objective, DoPTA consistently outperforms larger models, while using significantly lesser pre-training compute. DoPTA also sets new state-of-the art results on D4LA, and FUNSD, two challenging document visual analysis benchmarks.

文档理解图文对齐多模态零OCR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。