arXiv:2510.21774cs.CVcs.AI2025-10

构建首个真实场景的OCR质量评估数据集,支持模型训练与评测。

OCR-Quality: A Human-Annotated Dataset for OCR Quality Assessment

  • 1000页真实文档转为300DPI图片,涵盖多类型文本。
  • 用视觉语言模型预处理后人工打分,4级质量标准更精准。
  • 适合做OCR质量检测、验证系统开发的研究者使用。

我们提出OCR-Quality,一个用于评估和开发OCR质量评估方法的全面人工标注数据集。该数据集包含1000张从学术论文、教科书、电子书及多语言文档中采样的PDF页面转换成的300 DPI PNG图像,覆盖多样真实场景。每份文档经先进视觉-语言模型(VLMs)处理后,由人工依据四级评分体系(1:优秀,2:良好,3:一般,4:差)标注质量得分。数据集提供详细来源信息、标注指南及多种难度代表性案例。OCR-Quality填补了真实应用场景下可靠OCR质量评估的空白,为训练和评估OCR验证系统提供了重要基准。数据集已公开于https://huggingface.co/datasets/Aslan-mingye/OCR-Quality。

原文摘要 · Abstract (English)

We present OCR-Quality, a comprehensive human-annotated dataset designed for evaluating and developing OCR quality assessment methods. The dataset consists of 1,000 PDF pages converted to PNG images at 300 DPI, sampled from diverse real-world scenarios, including academic papers, textbooks, e-books, and multilingual documents. Each document has been processed using state-of-the-art Vision-Language Models (VLMs) and manually annotated with quality scores using a 4-level scoring system (1: Excellent, 2: Good, 3: Fair, 4: Poor). The dataset includes detailed source information, annotation guidelines, and representative cases across various difficulty levels. OCR-Quality addresses the critical need for reliable OCR quality assessment in real-world applications and provides a valuable benchmark for training and evaluating OCR verification systems. The dataset is publicly available at https://huggingface.co/datasets/Aslan-mingye/OCR-Quality .

OCR评估数据集视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。