构建首个评估大视觉语言模型少样本信息提取的基准数据集
CM1 -- A Dataset for Evaluating Few-Shot Information Extraction with Large Vision Language Models
- 针对手写文档设计少样本评测数据集,聚焦姓名与出生日期抽取
- 小样本下大模型性能超越传统全页模型,验证其泛化优势
- 适合研究少样本学习、文档智能与大模型应用的学者参考
从手写文档中自动提取关键信息是文档分析中的核心挑战,也是档案大规模数字化的前提。大型视觉语言模型(LVLM)在标注数据稀缺的场景下展现出巨大潜力。本文提出一个专为评估LVLM少样本能力而设计的新数据集——CM1,包含二战后欧洲为管理照护与维持计划所使用的手写表单。该数据集建立三个基准任务,用于提取姓名和出生日期,并涵盖不同训练集规模。我们为两种主流LVLM提供了基线结果,并与一个成熟的全页信息提取模型进行对比。尽管传统全页模型表现优异,但在仅少量训练样本时,所考察的LVLM凭借其庞大的参数量和预训练知识,显著优于传统方法。
原文摘要 · Abstract (English)
The automatic extraction of key-value information from handwritten documents is a key challenge in document analysis. A reliable extraction is a prerequisite for the mass digitization efforts of many archives. Large Vision Language Models (LVLM) are a promising technology to tackle this problem especially in scenarios where little annotated training data is available. In this work, we present a novel dataset specifically designed to evaluate the few-shot capabilities of LVLMs. The CM1 documents are a historic collection of forms with handwritten entries created in Europe to administer the Care and Maintenance program after World War Two. The dataset establishes three benchmarks on extracting name and birthdate information and, furthermore, considers different training set sizes. We provide baseline results for two different LVLMs and compare performances to an established full-page extraction model. While the traditional full-page model achieves highly competitive performances, our experiments show that when only a few training samples are available the considered LVLMs benefit from their size and heavy pretraining and outperform the classical approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。