用Transformer模型提升古籍手写文字识别准确率
Handwritten Text Recognition of Historical Manuscripts Using Transformer-Based Models
- 针对古籍手写特征设计四种新型数据增强方法
- 单模型CER达1.86,集成模型最优达1.60
- 适合历史文献数字化与数字人文研究者
历史手写文本识别(HTR)对挖掘档案文献的文化与学术价值至关重要,但常受限于标注数据稀少、语言变异及书写风格多样。本研究将最先进的基于Transformer的TrOCR模型应用于鲁道夫·格瓦尔特16世纪拉丁文手稿。我们探索了针对性图像预处理和广泛的数据增强技术,提出四种专为古籍手写特征设计的新增强方法,并评估集成学习以融合不同增强模型的互补优势。在格瓦尔特数据集上,最佳单一模型(弹性增强)达到1.86的字符错误率(CER),而前五名投票集成模型达到1.60的CER——相比最佳报告的TrOCR_BASE结果降低50%,较此前最优结果提升42%。结果表明,领域特定增强与集成策略显著提升了历史手稿的HTR性能。
原文摘要 · Abstract (English)
Historical handwritten text recognition (HTR) is essential for unlocking the cultural and scholarly value of archival documents, yet digitization is often hindered by scarce transcriptions, linguistic variation, and highly diverse handwriting styles. In this study, we apply TrOCR, a state-of-the-art transformer-based HTR model, to 16th-century Latin manuscripts authored by Rudolf Gwalther. We investigate targeted image preprocessing and a broad suite of data augmentation techniques, introducing four novel augmentation methods designed specifically for historical handwriting characteristics. We also evaluate ensemble learning approaches to leverage the complementary strengths of augmentation-trained models. On the Gwalther dataset, our best single-model augmentation (Elastic) achieves a Character Error Rate (CER) of 1.86, while a top-5 voting ensemble achieves a CER of 1.60 - representing a 50% relative improvement over the best reported TrOCR_BASE result and a 42% improvement over the previous state of the art. These results highlight the impact of domain-specific augmentations and ensemble strategies in advancing HTR performance for historical manuscripts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。