构建34万张波斯语合成图文对,解决波斯文识别数据匮乏问题。
Persian Pixel: A large-scale synthetic OCR dataset for Persian language

- 用700万词语料生成高保真图像,模拟波斯文连笔、变体等书写特征。
- 包含超34万对图文,覆盖句子到整页布局,支持端到端模型训练。
- 通过25种退化模型增强真实感,适合低资源复杂文字识别研究。
波斯语光学字符识别(OCR)发展远落后于拉丁语系,尽管全球有超过1.1亿人使用波斯语。这一差距源于波斯-阿拉伯文字系统的内在复杂性及高质量标注数据稀缺。波斯文具有强制连笔、上下文相关的字形变化、大量合字、变音符号位置及多种书写风格(如Naskh和Nastaliq),显著增加识别难度。同时,人工标注成本高、耗时长,导致数据瓶颈长期存在。本文提出Persian Pixel,一个大规模合成波斯语OCR数据集,包含超过34.3万对高质量图像与文本,基于七百万词的波斯语语料库,采用SynthOCR-Gen渲染框架生成,精准还原波斯文的排版特征,包括字形连接、位置变体、变音符号和多种主流字体。为缩小合成数据与真实场景差距,图像进一步添加25种随机退化模型,模拟墨水渗透、纸张老化、模糊、光照变化、扫描缺陷、压缩伪影及多种噪声。该数据集可直接用于训练和微调TrOCR、Donut等基于Transformer的OCR模型,为波斯语文档分析、历史手稿数字化和端到端文档理解提供坚实基础,证明程序化合成数据是推进低资源、复杂文字识别的有效低成本方案。
原文摘要 · Abstract (English)
Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages despite Persian being spoken by more than 110 million people across multiple countries. This gap arises from two fundamental challenges: the intrinsic complexity of the Perso-Arabic writing system and the limited availability of large-scale, high-quality annotated datasets. Persian script exhibits obligatory cursive connectivity, context-dependent glyph shaping, extensive ligatures, diacritic placement, and stylistic variation across writing forms such as Naskh and Nastaliq, all of which significantly complicate text recognition. At the same time, the high cost and labor-intensive nature of manual annotation have created a persistent data bottleneck, limiting the development of robust OCR systems and slowing progress in Persian document digitization.In this paper, we introduce Persian Pixel, a comprehensive synthetic OCR dataset specifically designed to address these challenges. Comprising over 343,000 high-fidelity image text pairs, the dataset spans sentence, paragraph, and full-page document layouts generated from a carefully curated seven-million-word Persian corpus using the SynthOCR-Gen rendering framework. The generation pipeline faithfully models the typographic characteristics of Persian script, including contextual character joining, positional glyph variants, diacritic placement, and multiple representative Persian typefaces. To bridge the synthetic-to-real domain gap, the rendered images are further enriched with more than twenty-five stochastic degradation models that emulate realistic document acquisition artifacts, including ink bleed, paper aging, blur, illumination variation, scanner imperfections, compression artifacts, and multiple noise processes.By overcoming the long-standing scarcity of annotated Persian OCR data, Persian Pixel provides a scalable and openly available resource for training and fine-tuning modern OCR architectures, including transformer-based models such as TrOCR and Donut. The dataset establishes a strong foundation for research in Persian document analysis, historical manuscript digitization, and end-to-end document understanding, while demonstrating that programmatic synthetic data generation offers a practical, cost-effective, and scalable alternative to manual annotation for advancing OCR in low-resource and typographically complex scripts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。