arXiv:2607.03650cs.CVcs.AI2026-07

公开临床扫描文档数据集,专为评估OCR模型设计

ClinOCR-Bench: A Comprehensive Clinical Scanned Document Dataset for Optical Character Recognition Model Evaluation

  • 构建包含6类扫描缺陷的384张真实医疗文档图像
  • 覆盖常见扫描瑕疵,支持系统性模型评估
  • 适合医疗AI研究者与医学影像处理团队使用

从外部实验室报告和手写表单等扫描医疗文档中提取文本信息,是现代电子健康记录(EHR)中的重大挑战。近年来,视觉语言模型(VLMs)在传统OCR工具之上展现出巨大潜力。然而,目前大多数临床OCR研究基于私有机构数据,公开可用的临床领域OCR评估数据集极少。此外,常见扫描伪影未被现有数据集充分反映,导致系统性评估难以开展。为此,我们发布了一个公开、逼真的OCR基准数据集ClinOCR-Bench,包含384张扫描图像,涵盖6个子集:正常、手写、低质量、旋转、表格和混合伪影。该数据集具备五大特点:1)多样化的文档类型与版式;2)全面覆盖常见EHR扫描伪影;3)无患者隐私信息;4)模板感知的训练/测试划分;5)充足的样本量以支持OCR基准测试。采用最先进的开源与专有VLM对基线模型性能进行了评估。数据集及文档已发布于GitHub(https://github.com/ClinOCR-Bench/ClinOCR-Bench)。

原文摘要 · Abstract (English)

Extracting textual information from scanned medical documents, such as external laboratory reports and manually filled forms, has been a major challenge in modern electronic health records (EHRs). Recent advancements in vision language models (VLMs) have shown great promise over traditional OCR tools. However, at this point, most clinical OCR studies were conducted on private, institutional data. To our knowledge, there are few publicly available datasets for evaluating OCR models in the clinical domain. Furthermore, common scanning artifacts that undermine OCR performance are not reflected in those datasets, leaving a systematic evaluation unfeasible. Therefore, we release a publicly available, realistic-looking OCR benchmark dataset, ClinOCR-Bench, with 384 scanned images across 6 subsets: Normal, Handwriting, Poor Quality, Rotation, Tables, and Mix-artifacts. ClinOCR-Bench features: 1) diverse document types and layouts, 2) full coverage of common EHR scan artifacts, 3) protected health information-free, 4) template-aware train/test split, and 5) adequate sample size for OCR benchmarking. Baseline OCR performance was evaluated using state-of-the-art open-weight and proprietary VLMs. The dataset and documentation are available on GitHub (https://github.com/ClinOCR-Bench/ClinOCR-Bench).

OCR医疗文本数据集视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。