arXiv:2606.06242cs.CLcs.AI2026-06

评测开源模型在机构文档中提取可复用数据快照的能力

Benchmarking Open-Source Layout Detection Models for Data Snapshot Extraction from Institutional Documents

论文配图:Benchmarking Open-Source Layout Detection Models for Data Snapshot Extraction from Institutional Documents
图 1 · 摘自论文原文
  • 构建专用数据集与评估框架,聚焦可复用分析图表的定位
  • 现有模型在机构文档上表现差,准确率低于60%且常误检非分析内容
  • 适合关注文档智能解析与政策研究自动化的研究者

机构文档中包含大量嵌入图表和表格的操作与分析信息。当前视觉内容提取方法多基于通用文档版面分析,将图表一视同仁处理,忽视其作为可复用分析实体的语义意义。本文提出「数据快照提取」任务,旨在识别并定位机构文档中具有分析价值的可视化内容。构建涵盖人道主义报告、世界银行政策研究报告及项目评估文件的基准数据集,标注了含可复用分析信息的图表。通过该数据集评测多个开源版面检测模型,结果表明:尽管在通用学术基准上表现良好,当前模型在实际机构文档中泛化能力差,常见错误包括混淆分析性与非分析性内容、复合分析图的碎片化分割、上下文信息提取不全。这些发现揭示了通用版面分析与实用数据快照提取间的显著差距。数据集与代码已开源,地址为https://huggingface.co/datasets/ai4data/data-snapshot及https://github.com/worldbank/ai4data/tree/main/experimental/data-snapshot。

原文摘要 · Abstract (English)

Institutional documents contain substantial amounts of operational and analytical information embedded within figures and tables. Current approaches for extracting visual content from documents are largely built around generic document layout analysis, where figures and tables are treated as uniformly relevant document objects rather than semantically meaningful analytical artifacts. In this work, we introduce a benchmark dataset and evaluation framework for \textit{data snapshot extraction}, the task of identifying and localizing semantically meaningful visual artifacts within institutional documents. The benchmark spans humanitarian reports, World Bank policy research working papers, and project appraisal documents, and includes annotations for figures and tables that contain reusable analytical information. Using this dataset, we benchmarked multiple open-source layout detection models and evaluated both detection performance and spatial extraction quality. Our results show that current models struggle to generalize to operational institutional documents despite strong performance on conventional academic benchmarks. Common failure modes include confusion between analytical and non-analytical content, fragmentation of composite analytical artifacts, and incomplete extraction of contextual information required for interpretation. These findings highlight a persistent gap between generic document layout analysis and operationally useful data snapshot extraction. We release the source PDFs, annotation dataset, metadata, and source code to support future research in operational document intelligence. The dataset is available at https://huggingface.co/datasets/ai4data/data-snapshot and the source code is available at https://github.com/worldbank/ai4data/tree/main/experimental/data-snapshot.

文档解析数据提取机构文档版面分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。