用跨语言联合训练提升阿拉伯字母手写文字识别在低资源下的表现
Cross-Lingual Learning within Arabic Script for Low-Resource HTR

- 利用共享书写系统的阿拉伯字母语言进行联合训练,缓解数据不足问题
- 在仅100-1000样本下,波斯语字符错误率降至9.99,乌尔都语降为14.45
- 适合研究低资源手写识别或跨语言迁移的开发者和学者
手写文本识别(HTR)在标注数据有限时仍具挑战性,尤其对阿拉伯字母语言。尽管序列模型在高资源场景表现良好,但数据稀缺时性能急剧下降。由于阿拉伯字母语言共享书写系统且字符重叠度高,跨语言学习可有效缓解数据短缺。本研究在低资源条件下(样本数K=100, 500, 1000)对阿拉伯(KHATT)、乌尔都(NUST-UHWR)和波斯(PHTD)三种语言进行线级联合训练。采用CRNN与基于视觉变换器的HTR-VT模型,在多个相关数据集联合训练后评估各目标语言性能。两种架构在低资源下均受益于跨语言训练:当目标语言数据极少时,CRNN更优;而随着目标语言数据增多,HTR-VT收益趋于不稳定。在波斯语(PHTD)上,联合训练达字符错误率(CER)9.99,超越此前结果;在额外乌尔都语数据集(UNHD)上,CER从17.20降至14.45。
原文摘要 · Abstract (English)
Handwritten Text Recognition (HTR) with limited labeled data remains a challenging problem, particularly for Arabic-script languages. Although modern sequence-based recognizers perform well in high-resource settings, their accuracy degrades sharply as training data becomes scarce. Arabic-script languages share a common writing system with substantial character overlap, motivating cross-lingual learning as a strategy to mitigate data scarcity. We conduct a controlled line-level study of cross-lingual joint training for Arabic-script HTR under low-resource regimes (number of samples K = 100, 500, 1000 labeled lines) on Arabic (KHATT), Urdu (NUST-UHWR) and Persian (PHTD). CRNN and Vision Transformer-based HTR-VT models are trained on the union of multiple related Arabic-script datasets to mitigate the data scarcity and are evaluated on individual target languages. Both architectures benefit from cross-language training under low-resource conditions. CRNN remains more effective under extremely limited target-language data, whereas the benefits of cross-language training for HTR-VT become less consistent as larger amounts of target-language data become available. On Persian (PHTD), joint training achieves a Character Error Rate (CER) of 9.99 , surpassing previously reported results despite not using the full available training data. On an additional Urdu dataset (UNHD), joint training reduces CER from 17.20 to 14.45.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。