用关系建模文档布局阅读顺序,提升视觉文档理解性能。
Modeling Layout Reading Order as Ordering Relations for Visually-rich Document Understanding

- 将布局阅读顺序建模为元素间的关系而非序列
- 新方法在两个任务设置上达当前最优效果
- 可通用增强任意视觉文档任务,适合做文档理解研究者
在视觉丰富文档(VrDs)中建模和利用布局阅读顺序对文档智能至关重要,因为它捕捉了文档内的丰富结构语义。以往工作通常将阅读顺序形式化为布局元素的排列,即包含所有元素的序列。然而我们认为这种形式无法充分表达布局中的完整阅读顺序信息,可能影响下游任务性能。为此,我们提出将阅读顺序建模为布局元素间的顺序关系,具备更强的表达能力。为支持该新形式的阅读顺序预测(ROP)方法评估,我们构建了一个综合性基准数据集,包含元素间阅读顺序关系标注,并提出一种基于关系抽取的方法,优于先前方法。此外,为凸显改进形式的实际效益,我们设计了一个阅读顺序关系增强管道,通过引入额外的阅读顺序关系输入,提升任意视觉文档任务模型性能。大量实验表明:(1)使用阅读顺序关系信息后,下游模型在目标数据集的两个任务设置上达到最优结果;(2)使用本模型生成的伪阅读顺序信息,无需针对性优化,在三种模型、八个跨领域文档信息抽取/问答任务设置中均实现性能提升。
原文摘要 · Abstract (English)
Modeling and leveraging layout reading order in visually-rich documents (VrDs) is critical in document intelligence as it captures the rich structure semantics within documents. Previous works typically formulated layout reading order as a permutation of layout elements, i.e. a sequence containing all the layout elements. However, we argue that this formulation does not adequately convey the complete reading order information in the layout, which may potentially lead to performance decline in downstream VrD tasks. To address this issue, we propose to model the layout reading order as ordering relations over the set of layout elements, which have sufficient expressive capability for the complete reading order information. To enable empirical evaluation on methods towards the improved form of reading order prediction (ROP), we establish a comprehensive benchmark dataset including the reading order annotation as relations over layout elements, together with a relation-extraction-based method that outperforms previous methods. Moreover, to highlight the practical benefits of introducing the improved form of layout reading order, we propose a reading-order-relation-enhancing pipeline to improve model performance on any arbitrary VrD task by introducing additional reading order relation inputs. Comprehensive results demonstrate that the pipeline generally benefits downstream VrD tasks: (1) with utilizing the reading order relation information, the enhanced downstream models achieve SOTA results on both two task settings of the targeted dataset; (2) with utilizing the pseudo reading order information generated by the proposed ROP model, the performance of the enhanced models has improved across all three models and eight cross-domain VrD-IE/QA task settings without targeted optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。