构建首个多文种哈萨克语合成OCR基准,揭示大模型在低资源脚本上的严重不足。
KazakhOCR: A Synthetic Benchmark for Evaluating Multimodal Models in Low-Resource Kazakh Script OCR
- 用7219张合成图像覆盖阿拉伯、西里尔、拉丁三套哈萨克语字体与噪声变化
- 三大主流多模态模型在拉丁/阿拉伯文识别中全失败,误判为阿拉伯语等其他语言
- 对比传统OCR发现大模型虽能理解语义但字符错误率远高于经典方法
哈萨克语是使用阿拉伯、西里尔和拉丁三种文字的突厥语族语言,在光学字符识别(OCR)方面具有独特性。针对低资源哈萨克语文字的OCR研究极少,且缺乏阿拉伯文和拉丁文的基准数据集与图像。本文构建了一个包含7,219张图像的合成OCR数据集,涵盖三种文字,并引入字体、颜色和噪声变化以模拟真实场景。我们评估了三个多模态大语言模型(MLLMs):Gemma-3-12B-it、Qwen2.5-VL-7B-Instruct 和 Llama-3.2-11B-Vision-Instruct,发现它们在拉丁文和阿拉伯文识别上均失败,且无法将阿拉伯文正确识别为哈萨克语,误分类为阿拉伯语、波斯语和库尔德语。进一步与经典OCR基线比较表明,虽然传统方法字符错误率更低,但大模型在性能上仍难以匹敌。结果揭示当前多模态大模型处理低资源音节型文字的能力存在显著差距,凸显开发包容性模型与支持低资源语言的基准的必要性。
原文摘要 · Abstract (English)
Kazakh is a Turkic language using the Arabic, Cyrillic, and Latin scripts, making it unique in terms of optical character recognition (OCR). Work on OCR for low-resource Kazakh scripts is very scarce, and no OCR benchmarks or images exist for the Arabic and Latin scripts. We construct a synthetic OCR dataset of 7,219 images for all three scripts with font, color, and noise variations to imitate real OCR tasks. We evaluated three multimodal large language models (MLLMs) on a subset of the benchmark for OCR and language identification: Gemma-3-12B-it, Qwen2.5-VL-7B-Instruct, and Llama-3.2-11B-Vision-Instruct. All models are unsuccessful with Latin and Arabic script OCR, and fail to recognize the Arabic script as Kazakh text, misclassifying it as Arabic, Farsi, and Kurdish. We further compare MLLMs with a classical OCR baseline and find that while traditional OCR has lower character error rates, MLLMs fail to match this performance. These findings show significant gaps in current MLLM capabilities to process low-resource Abjad-based scripts and demonstrate the need for inclusive models and benchmarks supporting low-resource scripts and languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。