用图结构提升文档多模态信息提取效果
GraphRevisedIE: Multimodal Information Extraction with Graph-Revised Network
- 构建图网络融合文本、视觉与布局特征
- 在多个数据集上表现优于或接近主流方法
- 适合研究文档智能与多模态提取的开发者
从视觉丰富文档(VRD)中提取关键信息是文档智能中的挑战任务,主要源于文档版式复杂多样导致模型泛化困难,以及缺乏有效利用多模态特征的方法。本文提出轻量级模型 GraphRevisedIE,能够有效融合文本、视觉和布局等多模态特征,并通过图修订与图卷积引入全局上下文信息以增强嵌入表示。在多个真实世界数据集上的大量实验表明,GraphRevisedIE 能够适应不同版式的文档,性能达到或超过现有 KIE 方法。此外,我们发布了一个包含真实与合成样本的企业营业执照数据集,以推动文档 KIE 研究的发展。
原文摘要 · Abstract (English)
Key information extraction (KIE) from visually rich documents (VRD) has been a challenging task in document intelligence because of not only the complicated and diverse layouts of VRD that make the model hard to generalize but also the lack of methods to exploit the multimodal features in VRD. In this paper, we propose a light-weight model named GraphRevisedIE that effectively embeds multimodal features such as textual, visual, and layout features from VRD and leverages graph revision and graph convolution to enrich the multimodal embedding with global context. Extensive experiments on multiple real-world datasets show that GraphRevisedIE generalizes to documents of varied layouts and achieves comparable or better performance compared to previous KIE methods. We also publish a business license dataset that contains both real-life and synthesized documents to facilitate research of document KIE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。