arXiv:2501.03403cs.CLcs.AI2025-01被引 10

构建首个含空间标注的文档问答数据集,助力大模型理解图文内容。

BoundingDocs: a Unified Dataset for Document Question Answering with Spatial Annotations

  • 将信息抽取等任务转为问答形式,适配大模型训练与评估
  • 提供所有文档的OCR文本及答案在图像中的精确位置框
  • 验证带空间信息的提示词对模型理解能力的提升效果

我们提出一个统一的文档问答(Document QA)数据集,整合了多个与文档AI和视觉丰富文档理解(VRDU)相关的公开数据集。主要贡献有二:一方面,将现有文档AI任务(如信息抽取)重构为问答任务,使其适合用于大语言模型的训练与评估;另一方面,公开所有文档的OCR文本,并标注答案在文档图像中的精确位置(边界框)。利用该数据集,我们研究了不同提示技术(可能包含边界框信息)对开源模型性能的影响,识别出最有效的文档理解方法。

原文摘要 · Abstract (English)

We present a unified dataset for document Question-Answering (QA), which is obtained combining several public datasets related to Document AI and visually rich document understanding (VRDU). Our main contribution is twofold: on the one hand we reformulate existing Document AI tasks, such as Information Extraction (IE), into a Question-Answering task, making it a suitable resource for training and evaluating Large Language Models; on the other hand, we release the OCR of all the documents and include the exact position of the answer to be found in the document image as a bounding box. Using this dataset, we explore the impact of different prompting techniques (that might include bounding box information) on the performance of open-weight models, identifying the most effective approaches for document comprehension.

文档理解问答系统空间标注大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。