arXiv:2409.00447cs.AI2024-09被引 4

构建可解释成绩单数据集,用于训练和测试视觉文档理解模型。

The MERIT Dataset: Modelling and Efficiently Rendering Interpretable Transcripts

  • 基于学生成绩单构建多模态标注数据集,含文本、图像与版式信息。
  • 包含33,000样本与400+标签,挑战当前顶级模型的文本识别能力。
  • 支持评估大模型在真实场景中的偏见问题,适合文档理解研究者使用。

本文提出MERIT数据集,一个以学生成绩单为背景的多模态(文本+图像+版式)全标注数据集。该数据集包含超过400个标签和33,000个样本,是开展复杂视觉丰富文档理解(VrDU)任务建模的重要资源。由于其内容特性(学生成绩报告),该数据集可可控地引入偏见,因而成为评估语言模型(LLMs)中偏见传播的有效工具。论文详细描述了数据集的生成流程,并重点阐述其在文本、视觉、版式及偏见维度上的特征。为验证数据集实用性,我们设计了基于词元分类的基准测试,结果表明即便对最先进模型而言,该数据集仍构成显著挑战,且通过将其样本纳入预训练可显著提升模型性能。

原文摘要 · Abstract (English)

This paper introduces the MERIT Dataset, a multimodal (text + image + layout) fully labeled dataset within the context of school reports. Comprising over 400 labels and 33k samples, the MERIT Dataset is a valuable resource for training models in demanding Visually-rich Document Understanding (VrDU) tasks. By its nature (student grade reports), the MERIT Dataset can potentially include biases in a controlled way, making it a valuable tool to benchmark biases induced in Language Models (LLMs). The paper outlines the dataset's generation pipeline and highlights its main features in the textual, visual, layout, and bias domains. To demonstrate the dataset's utility, we present a benchmark with token classification models, showing that the dataset poses a significant challenge even for SOTA models and that these would greatly benefit from including samples from the MERIT Dataset in their pretraining phase.

文档理解多模态偏见评估数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。