arXiv:2409.15004cs.AIcs.CL2024-09中稿 · MIDAS被引 2

改进模型从无结构金融文档中精准提取关键信息

ViBERTgrid BiLSTM-CRF: Multimodal Key Information Extraction from Unstructured Financial Documents

  • 在ViBERTgrid基础上加入BiLSTM-CRF层处理无结构文本
  • 金融领域命名实体识别准确率提升最高达2个百分点
  • 公开了SROIE数据集的细粒度标注,利于后续研究

多模态关键信息提取(KIE)模型在半结构化文档上已有广泛研究,但在无结构文档上的探索仍属新兴方向。本文提出将先前用于半结构化文档的多模态Transformer模型ViBERTgrid适配至无结构金融文档,通过引入BiLSTM-CRF层实现改进。所提出的ViBERTgrid BiLSTM-CRF模型在金融领域无结构文档的命名实体识别任务上表现显著提升,性能最高提升2个百分点,同时保持了在半结构化文档上的KIE性能。此外,本文公开发布了SROIE数据集的逐标记级标注,以促进多模态序列标注模型的研究与应用。

原文摘要 · Abstract (English)

Multimodal key information extraction (KIE) models have been studied extensively on semi-structured documents. However, their investigation on unstructured documents is an emerging research topic. The paper presents an approach to adapt a multimodal transformer (i.e., ViBERTgrid previously explored on semi-structured documents) for unstructured financial documents, by incorporating a BiLSTM-CRF layer. The proposed ViBERTgrid BiLSTM-CRF model demonstrates a significant improvement in performance (up to 2 percentage points) on named entity recognition from unstructured documents in financial domain, while maintaining its KIE performance on semi-structured documents. As an additional contribution, we publicly released token-level annotations for the SROIE dataset in order to pave the way for its use in multimodal sequence labeling models.

多模态信息提取金融文本序列标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。