用生成式组合器解决少样本文档信息提取难题
Generative Compositor for Few-Shot Visual Information Extraction
- 通过混合指针-生成网络模拟排版操作,按提示词组合文本
- 1-shot到10-shot设置下显著优于基线模型
- 适合低资源场景下的文档结构化信息抽取
视觉信息提取(VIE)旨在从视觉丰富的文档图像中提取结构化信息,在文档处理中具有关键作用。由于布局、语义范围和语言的多样性,VIE可能包含数千种类型,但许多类型缺乏训练数据,带来巨大挑战。本文提出一种新型生成模型——生成式组合器(Generative Compositor),用于解决少样本VIE问题。该模型是一种混合指针-生成网络,通过检索源文本中的词语并根据提示词进行组装,模拟排版操作。此外,采用三种预训练策略以增强模型对空间上下文信息的感知能力。同时设计了提示词感知重采样器,利用提示词中的实体语义先验实现高效匹配。基于提示的检索机制与预训练策略使模型在少量训练样本下仍能获得有效空间与语义线索。实验表明,该方法在全量样本训练下表现优异,而在1-shot、5-shot和10-shot设置下显著超越基线模型。
原文摘要 · Abstract (English)
Visual Information Extraction (VIE), aiming at extracting structured information from visually rich document images, plays a pivotal role in document processing. Considering various layouts, semantic scopes, and languages, VIE encompasses an extensive range of types, potentially numbering in the thousands. However, many of these types suffer from a lack of training data, which poses significant challenges. In this paper, we propose a novel generative model, named Generative Compositor, to address the challenge of few-shot VIE. The Generative Compositor is a hybrid pointer-generator network that emulates the operations of a compositor by retrieving words from the source text and assembling them based on the provided prompts. Furthermore, three pre-training strategies are employed to enhance the model's perception of spatial context information. Besides, a prompt-aware resampler is specially designed to enable efficient matching by leveraging the entity-semantic prior contained in prompts. The introduction of the prompt-based retrieval mechanism and the pre-training strategies enable the model to acquire more effective spatial and semantic clues with limited training samples. Experiments demonstrate that the proposed method achieves highly competitive results in the full-sample training, while notably outperforms the baseline in the 1-shot, 5-shot, and 10-shot settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。