arXiv:2606.18884cs.CV2026-06中稿 · TIPS workshop ICPR…

对比阿拉伯语与拉丁语手写识别,发现差距源于字符相似性和数据量不足。

Performance Gap Analysis between Latin and Arabic Scripts HTR

论文配图:Performance Gap Analysis between Latin and Arabic Scripts HTR
图 1 · 摘自论文原文
  • 统一使用CRNN模型在9个数据集上对比,控制变量分析性能差异
  • 阿拉伯语识别错误率始终高5-7个字符错误率点,即使全量数据仍存差距
  • 字符视觉相似性高、分布更偏斜,适合研究低资源场景与标注质量改进

近期研究表明,手写文本识别(HTR)系统在阿拉伯语数据集上的表现逊于拉丁语。然而,由于缺乏受控比较,这一差距的原因仍不明确。本文使用统一的CRNN模型,在九个数据集(包括KHATT、Muharaf、NUST-UHWR、PHTD、IAM、READ-2016等)和不同训练规模(K ∈ {100, 500, 1000, 2000, ..., Kfull})下进行线级手写识别的全面对比。结果表明,性能差距持续存在:在低资源情况下显著,随数据增加而缩小,但在全量数据下仍保持5-7个字符错误率(CER)点的差距。我们发现标注质量至关重要,多个数据集包含标签错误;清洗后可降低错误率并缩小差距,但无法消除。此外,固定数量样本在阿拉伯语中覆盖效果较差,因视觉变异性更高,需更多数据学习一致表征。通过比较文本行数与字符数,揭示其等价权衡关系。分析字符频率分布发现,阿拉伯语显著更重尾,且约30%的替换错误源于视觉相似字符混淆,远高于拉丁语(约15%),如在KHATT与IAM数据集中的对比。

原文摘要 · Abstract (English)

Recent studies have shown that handwritten text recognition (HTR) systems perform worse on Arabic-script datasets than on Latin-script data. However, the reasons for this gap are still not well understood due to the lack of controlled comparisons. In this work, we present a comprehensive study of Arabic and Latin scripts HTR using a unified CRNN model for line-level HTR across nine datasets (including KHATT (Arabic), Muharaf (Arabic), NUST-UHWR (Urdu), PHTD (Persian), IAM (English), READ-2016 (German), and others) and di ferent training sizes (K in {100, 500, 1000, 2000, ..., Kfull}). Our results show the performance gap remains: it is large in low-resource settings, decreases with more data, but remains even at full scale, with a consistent difference of 5-7 CER points. We show that annotation quality matters, as many datasets contain labeling errors. Cleaning reduces error rates and narrows the gap, but does not eliminate it. In addition, we find that a fixed number of training samples provides less effective coverage in Arabic due to higher visual variability, requiring more data to learn similar representations. We compare recognition across datasets in terms of the number of text lines and the number of characters, showing an equivalence trade-off. We compare character frequency distributions across scripts and show that Arabic is significantly more heavy-tailed than Latin. Our error analysis reveals that around 30 percent of substitution errors in Arabic datasets (e.g., KHATT) are caused by confusion between visually similar characters, compared to about 15 percent in Latin-script datasets such as IAM.

手写识别阿拉伯语数据偏差字符混淆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。