首个针对土耳其语文档解析的综合性基准,覆盖多类文档与难度等级。
OCRTurk: A Comprehensive OCR Benchmark for Turkish
- 构建涵盖180份真实文档的土耳其语解析基准OCRTurk。
- PaddleOCR在多数元素识别上表现最优,难样本下仍保持高编辑距离得分。
- 滑灯片类文档最难,非学术文档更易处理,适合研究低资源语言场景。
文档解析广泛应用于大规模文档数字化、检索增强生成及医疗教育等专业流程。评估模型可靠性与实际鲁棒性需依赖基准测试。现有基准多面向高资源语言,对土耳其语等低资源语言覆盖不足,且缺乏反映真实场景与文档多样性的标准评测体系。为此,我们提出OCRTurk,一个涵盖多种版面元素与文档类别、分三个难度级别的土耳其语文档解析基准。OCRTurk包含180份来自学术论文、学位论文、幻灯片和非学术文章的真实文档。我们在该基准上评估了七种OCR模型,采用细粒度元素级指标。在各难度层级中,PaddleOCR整体表现最强,除图像外多数元素识别指标领先,并在简单、中等、困难子集上均取得高归一化编辑距离分数。此外,模型在非学术文档上表现良好,而幻灯片类文档最为挑战。
原文摘要 · Abstract (English)
Document parsing is now widely used in applications, such as large-scale document digitization, retrieval-augmented generation, and domain-specific pipelines in healthcare and education. Benchmarking these models is crucial for assessing their reliability and practical robustness. Existing benchmarks mostly target high-resource languages and provide limited coverage for low-resource settings, such as Turkish. Moreover, existing studies on Turkish document parsing lack a standardized benchmark that reflects real-world scenarios and document diversity. To address this gap, we introduce OCRTurk, a Turkish document parsing benchmark covering multiple layout elements and document categories at three difficulty levels. OCRTurk consists of 180 Turkish documents drawn from academic articles, theses, slide decks, and non-academic articles. We evaluate seven OCR models on OCRTurk using element-wise metrics. Across difficulty levels, PaddleOCR achieves the strongest overall results, leading most element-wise metrics except figures and attaining high Normalized Edit Distance scores in easy, medium, and hard subsets. We also observe performance variation by document type. Models perform well on non-academic documents, while slideshows become the most challenging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。