用多模态大模型提升阿拉伯文识别精度,支持连笔、符号和多种字体。
QARI-OCR: High-Fidelity Arabic Text Recognition through Multimodal Large Language Model Adaptation
- 基于Qwen2-VL迭代优化,专攻阿拉伯文复杂书写特征。
- 在带符号文本上达到WER 0.160、CER 0.061的顶尖水平。
- 适合需要高精度阿拉伯文识别的研究与应用开发者。
阿拉伯文书写固有的复杂性——连笔形式、元音符号(tashkeel)及多样字体——长期困扰光学字符识别(OCR)。本文提出Qari-OCR,一系列基于Qwen2-VL-2B-Instruct的视觉语言模型,通过在专用合成数据集上逐步微调实现对阿拉伯文的优化。其领先模型QARI v0.2在含丰富tashkeel的文本上达到词错误率(WER)0.160、字符错误率(CER)0.061、BLEU分数0.737,创下开源新纪录。Qari-OCR在处理元音符号、不同字体与文档布局方面表现优异,并在低分辨率图像上仍保持强鲁棒性。后续探索(QARI v0.3)展现出结构化文档理解与手写体识别的潜力。本研究显著提升了阿拉伯文OCR的准确率与效率,所有模型与数据集均已开源,以推动进一步研究。
原文摘要 · Abstract (English)
The inherent complexities of Arabic script; its cursive nature, diacritical marks (tashkeel), and varied typography, pose persistent challenges for Optical Character Recognition (OCR). We present Qari-OCR, a series of vision-language models derived from Qwen2-VL-2B-Instruct, progressively optimized for Arabic through iterative fine-tuning on specialized synthetic datasets. Our leading model, QARI v0.2, establishes a new open-source state-of-the-art with a Word Error Rate (WER) of 0.160, Character Error Rate (CER) of 0.061, and BLEU score of 0.737 on diacritically-rich texts. Qari-OCR demonstrates superior handling of tashkeel, diverse fonts, and document layouts, alongside impressive performance on low-resolution images. Further explorations (QARI v0.3) showcase strong potential for structural document understanding and handwritten text. This work delivers a marked improvement in Arabic OCR accuracy and efficiency, with all models and datasets released to foster further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。