arXiv:2410.10471cs.CV2024-10被引 7

提出新文档理解任务与方法,无需人工分组也能精准识别语义关系。

ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-training

  • 通过排列文字让同语义词表示更接近,自动学习语义分组。
  • 在无手动分组条件下,性能显著优于现有方法。
  • 适合真实场景下的文档理解,尤其对复杂排版有效。

当前视觉丰富文档理解(VrDU)方法依赖人工标注的语义分组,而这些分组无法通过OCR自动获取,导致实际应用受限。为此,我们提出真实世界视觉丰富文档理解(ReVrDU),不再使用人工标注的语义分组。同时,我们提出新方法ReLayout,通过排列文本并拉近同一语义组内词语的表示,实现语义分组的自学习。实验表明,传统方法在ReVrDU任务上性能明显下降,而ReLayout展现出更优表现。

原文摘要 · Abstract (English)

Recent approaches for visually-rich document understanding (VrDU) uses manually annotated semantic groups, where a semantic group encompasses all semantically relevant but not obviously grouped words. As OCR tools are unable to automatically identify such grouping, we argue that current VrDU approaches are unrealistic. We thus introduce a new variant of the VrDU task, real-world visually-rich document understanding (ReVrDU), that does not allow for using manually annotated semantic groups. We also propose a new method, ReLayout, compliant with the ReVrDU scenario, which learns to capture semantic grouping through arranging words and bringing the representations of words that belong to the potential same semantic group closer together. Our experimental results demonstrate the performance of existing methods is deteriorated with the ReVrDU task, while ReLayout shows superiour performance.

文档理解语义分组预训练布局建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。