arXiv:2503.16868cs.CLcs.CV2025-03

联合提取多字段信息,提升文档理解准确率

Joint Extraction Matters: Prompt-Based Visual Question Answering for Multi-Field Document Information Extraction

  • 用多字段联合提示词替代单字段独立提问
  • 联合提取在数值或上下文强关联时准确率更高
  • 适合需要跨字段推理的文档信息抽取场景

视觉问答(VQA)已成为从文档图像中提取特定信息的灵活方法。然而,现有工作通常孤立地查询每个字段,忽略了多个字段间的潜在依赖关系。本文研究了联合提取多个字段与单独提取的优劣。在多个大型视觉语言模型和数据集上的实验表明,联合提取能显著提升准确率,尤其当字段间存在强数值或上下文依赖时。我们进一步分析了请求字段数量对性能的影响,并采用基于回归的度量量化字段间关系。结果表明,多字段提示词可缓解表面形式相似或数值相近带来的混淆,为文档信息抽取中的VQA系统设计提供实用方法。

原文摘要 · Abstract (English)

Visual question answering (VQA) has emerged as a flexible approach for extracting specific pieces of information from document images. However, existing work typically queries each field in isolation, overlooking potential dependencies across multiple items. This paper investigates the merits of extracting multiple fields jointly versus separately. Through experiments on multiple large vision language models and datasets, we show that jointly extracting fields often improves accuracy, especially when the fields share strong numeric or contextual dependencies. We further analyze how performance scales with the number of requested items and use a regression based metric to quantify inter field relationships. Our results suggest that multi field prompts can mitigate confusion arising from similar surface forms and related numeric values, providing practical methods for designing robust VQA systems in document information extraction tasks.

视觉问答文档理解联合提取多字段

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。