arXiv:2410.15484cs.CL2024-10EMNLP被引 8

用多样模板提升文档信息抽取数据质量,让大模型更懂真实场景问题。

"What is the value of {templates}?" Rethinking Document Information Extraction Datasets for LLMs

  • 设计五套定制化提问模板,覆盖多实体、提取与判断类问题。
  • 在K2Q上测试的生成模型表现优于简单模板,零样本性能显著提升。
  • 适合研究高质量训练数据构建,尤其关注工业级文档理解的团队。

大语言模型在视觉丰富文档理解(VRDU)中的兴起,推动了基于提示-响应的文档数据集需求。由于从头标注成本高昂,现有研究多通过简单模板生成新数据集。针对关键信息抽取(KIE)这一常见任务,以往普遍使用“{key}的值是什么?”这类统一模板。然而,真实场景中问题形式多样,单一模板难以支撑鲁棒模型。本文提出K2Q,一套由KIE转化而来的多样化数据集,包含五组不同设计的提示模板,支持多实体、可提取或布尔判断型问题。我们对七种基线生成模型在K2Q上进行零样本评估,并对比其中三种模型在K2Q与简单模板上的训练表现。结果表明,使用复杂多样的提问方式能显著提升模型性能与鲁棒性。本工作旨在推动生成式模型训练数据质量的研究。

原文摘要 · Abstract (English)

The rise of large language models (LLMs) for visually rich document understanding (VRDU) has kindled a need for prompt-response, document-based datasets. As annotating new datasets from scratch is labor-intensive, the existing literature has generated prompt-response datasets from available resources using simple templates. For the case of key information extraction (KIE), one of the most common VRDU tasks, past work has typically employed the template "What is the value for the {key}?". However, given the variety of questions encountered in the wild, simple and uniform templates are insufficient for creating robust models in research and industrial contexts. In this work, we present K2Q, a diverse collection of five datasets converted from KIE to a prompt-response format using a plethora of bespoke templates. The questions in K2Q can span multiple entities and be extractive or boolean. We empirically compare the performance of seven baseline generative models on K2Q with zero-shot prompting. We further compare three of these models when training on K2Q versus training on simpler templates to motivate the need of our work. We find that creating diverse and intricate KIE questions enhances the performance and robustness of VRDU models. We hope this work encourages future studies on data quality for generative model training.

信息抽取提示工程数据质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。