arXiv:2601.11425cs.CVcs.CL2026-01

构建了百万页科学论文的精准文字定位数据集,支持版面分析与基于坐标的问答。

PubMed-OCR: PMC Open Access OCR Annotations

  • 从PMC开放论文中提取图像,用Google Vision标注文字层级边界框
  • 覆盖20.9万篇文章、150万页,含约13亿词,支持坐标感知任务
  • 适合做版面理解、OCR评估和基于位置的学术问答研究

PubMed-OCR是一个以OCR为核心的科学论文语料库,源自PubMed Central的开放获取PDF文档。每页图像均通过Google Cloud Vision进行标注,以紧凑的JSON格式发布,包含字、行、段落级别的边界框信息。该语料库涵盖20.9万篇论文(共150万页,约13亿词),可用于布局感知建模、坐标驱动的问答任务以及依赖OCR流水线的评估。我们分析了语料库特征(如期刊覆盖范围与检测到的版面元素),并讨论了局限性,包括仅依赖单一OCR引擎及启发式行重建方法。数据与模式已公开,以促进下游研究,并欢迎进一步扩展。

原文摘要 · Abstract (English)

PubMed-OCR is an OCR-centric corpus of scientific articles derived from PubMed Central Open Access PDFs. Each page image is annotated with Google Cloud Vision and released in a compact JSON schema with word-, line-, and paragraph-level bounding boxes. The corpus spans 209.5K articles (1.5M pages; ~1.3B words) and supports layout-aware modeling, coordinate-grounded QA, and evaluation of OCR-dependent pipelines. We analyze corpus characteristics (e.g., journal coverage and detected layout features) and discuss limitations, including reliance on a single OCR engine and heuristic line reconstruction. We release the data and schema to facilitate downstream research and invite extensions.

OCR科学文本版面分析数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。