arXiv:2410.06802cs.CL2024-10EMNLP被引 2

将文档结构解析为动作序列生成,提升长文档理解能力

Seg2Act: Global Context-aware Action Generation for Document Logical Structuring

  • 把文档结构提取看作动作序列生成任务,迭代更新上下文与结构
  • 在ChCatExt和HierDoc上表现优于传统方法,支持监督与迁移学习
  • 适合需要精准文档逻辑分析的智能系统开发者

文档逻辑结构化旨在提取文档的潜在层次结构,这对文档智能至关重要。传统方法难以应对长文档的复杂性与多样性。为此,我们提出Seg2Act,一种端到端的生成式文档逻辑结构化方法,将逻辑结构提取重新定义为动作生成任务。给定文档的文本片段,Seg2Act通过全局上下文感知生成模型迭代生成动作序列,并根据已生成动作同步更新全局上下文与当前逻辑结构。在ChCatExt和HierDoc数据集上的实验表明,Seg2Act在监督学习和迁移学习设置下均表现出色。

原文摘要 · Abstract (English)

Document logical structuring aims to extract the underlying hierarchical structure of documents, which is crucial for document intelligence. Traditional approaches often fall short in handling the complexity and the variability of lengthy documents. To address these issues, we introduce Seg2Act, an end-to-end, generation-based method for document logical structuring, revisiting logical structure extraction as an action generation task. Specifically, given the text segments of a document, Seg2Act iteratively generates the action sequence via a global context-aware generative model, and simultaneously updates its global context and current logical structure based on the generated actions. Experiments on ChCatExt and HierDoc datasets demonstrate the superior performance of Seg2Act in both supervised and transfer learning settings.

文档结构生成模型逻辑分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。