现代手写预训练对历史阿拉伯文识别效果不确定,实验证明无明显增益。
Does a Modern-Handwriting Warm-Up Help Historical Arabic OCR? A Reproducible, Compute-Matched Evaluation on Muharaf and KHATT

- 四次重复实验显示预训练效果正负波动大,仅两组干净实验表明无显著影响。
- 在计算资源匹配下,预训练反而使识别错误率升高2.42个点,但差异不显著。
- 开源完整工具链,支持结果复现,强调实验可重复性的重要性。
现有研究常基于单一实现判断现代阿拉伯手写预训练是否有助于历史阿拉伯文文本识别(HTR),证据不足。本文通过四次重复实验,改变基础检查点、编码器冻结策略、训练轮数、精度和学习率调度等自然变化因素,固定归一化、评分器与置信区间估计,发现预训练效果在-17.64至+14.52 CER点间剧烈波动并反转符号。其中两个极端结果由可识别混淆因素导致(学习率低五倍、检查点来源不明);其余两组清洁实验结果为-0.25与+0.94,即无显著影响。进一步进行计算量匹配实验,在三个随机种子下,使用相同领域控制组对比,结果显示预训练反而更差,平均恶化2.42 CER点(95%置信区间[+0.60, +4.25]);仅约0.6点来自书写风格差异,影响微小,非普遍结论。研究发布SaudiHeritage-OCR包,包含归一化模块、置信区间评分器、经验证的KHATT解码器、实验清单、视觉语言模型基线及版本对齐协议,确保结果可独立验证。阿尔-马赫德铭文保持严格隔离,不作为基准测试。
原文摘要 · Abstract (English)
Whether an intermediate stage of modern Arabic handwriting helps or hurts historical Arabic HTR is usually decided from one implementation and one comparison, too thin a basis for a claim either way. We test stability by running the same nominal ablation four times, letting the base checkpoint, encoder-freezing strategy, epoch budget, precision, and learning-rate schedule vary as they naturally did during development, while holding the normalization, scorer, and interval estimation fixed. Each run compares intermediate training on modern handwriting (KHATT) then fine-tuning on historical manuscripts (Muharaf) against fine-tuning on Muharaf directly. Across the four runs the estimated effect swings from -17.64 to +14.52 CER points and reverses sign. The two extremes are exactly the two runs with an identifiable confound (a fivefold lower learning rate in one; a checkpoint of undisclosed provenance in the other); the two clean runs land at -0.25 and +0.94, i.e. no effect. A tight interval from one implementation says nothing about the next. We then run a compute-matched experiment with identical budgets over three seeds: KHATT warm-up is +2.42 CER points worse than a matched same-domain control (95% interval [+0.60, +4.25]); the part of that gap specific to the handwriting domain is only about 0.6 points a small negative effect under this configuration, not a universal result. We release a SaudiHeritage-OCR package with the normalizer, interval scorer, a verified KHATT decoder, experimental manifests, VLM baselines, and an edition-alignment protocol, so the result can be checked independently. The Al-Mahd inscription line is held strictly out and is not offered as a benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。