构建大规模合成阿拉伯文书籍文本数据集,提升识别模型泛化能力。
SARD: A Large-Scale Synthetic Arabic OCR Dataset for Book-Style Text Recognition
- 通过合成生成84万张书籍布局图像,覆盖10种字体。
- 含6.9亿个单词,无真实噪声,适合训练高鲁棒性模型。
- 适合研究阿拉伯文版面识别、视觉语言模型的开发者使用。
阿拉伯文光学字符识别(OCR)对数字化大量阿拉伯印刷资料至关重要。然而,训练现代OCR模型,尤其是强大的视觉-语言模型,受限于缺乏大规模、多样化且结构逼真的真实书籍布局数据集。现有阿拉伯文OCR数据集多聚焦孤立单词或单行文本,规模有限,且在字体多样性与版式复杂度方面不足。为弥补这一空白,我们提出SARD(大规模合成阿拉伯文书籍文本识别数据集)。SARD是一个大规模、合成生成的数据集,专为模拟书籍文档设计,包含843,622张文档图像,共6.9亿个单词,覆盖10种不同阿拉伯字体,确保广泛的排版风格覆盖。与扫描文档数据集不同,SARD无真实世界噪声和失真,提供干净可控的训练环境。其合成特性具备极强可扩展性,并可精确控制版式与内容变化。我们详细说明了数据集构成与生成流程,并提供了多个OCR模型(包括传统与深度学习方法)的基准测试结果,揭示该数据集带来的挑战与机遇。SARD为开发和评估能够处理多样阿拉伯文书籍文本的鲁棒型OCR及视觉-语言模型提供了宝贵资源。
原文摘要 · Abstract (English)
Arabic Optical Character Recognition (OCR) is essential for converting vast amounts of Arabic print media into digital formats. However, training modern OCR models, especially powerful vision-language models, is hampered by the lack of large, diverse, and well-structured datasets that mimic real-world book layouts. Existing Arabic OCR datasets often focus on isolated words or lines or are limited in scale, typographic variety, or structural complexity found in books. To address this significant gap, we introduce SARD (Large-Scale Synthetic Arabic OCR Dataset). SARD is a massive, synthetically generated dataset specifically designed to simulate book-style documents. It comprises 843,622 document images containing 690 million words, rendered across ten distinct Arabic fonts to ensure broad typographic coverage. Unlike datasets derived from scanned documents, SARD is free from real-world noise and distortions, offering a clean and controlled environment for model training. Its synthetic nature provides unparalleled scalability and allows for precise control over layout and content variation. We detail the dataset's composition and generation process and provide benchmark results for several OCR models, including traditional and deep learning approaches, highlighting the challenges and opportunities presented by this dataset. SARD serves as a valuable resource for developing and evaluating robust OCR and vision-language models capable of processing diverse Arabic book-style texts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。