用历史与现代字体配对图像提升古日文识别准确率
Training Kindai OCR with parallel textline images and self-attention feature distance-based loss
- 用平行文本行图像增强训练数据,缓解古文标注数据少的问题
- 自注意力特征距离损失使识别错误率降低2.23%至3.94%
- 适合做历史文献数字化的NLP与计算机视觉研究者参考
明治时期文献(19世纪末至20世纪初)以现代日语书写,对研究当时社会结构、日常生活与环境具有重要历史价值。但其文字转录工作繁重,导致用于光学字符识别(OCR)训练的标注数据稀缺。本文通过使用原始明治文本与其对应现代日文字体的平行文本行图像,扩充训练数据集。提出一种基于自注意力特征距离的损失函数,最小化成对图像间特征差异,采用欧氏距离与最大均值差异(MMD)作为域适应度量。实验表明,该方法在基于Transformer的OCR基线基础上,分别将字符错误率(CER)降低2.23%和3.94%。此外,该方法提升了自注意力表示的判别能力,显著改善了历史文档的识别性能。
原文摘要 · Abstract (English)
Kindai documents, written in modern Japanese from the late 19th to early 20th century, hold significant historical value for researchers studying societal structures, daily life, and environmental conditions of that period. However, transcribing these documents remains a labor-intensive and time-consuming task, resulting in limited annotated data for training optical character recognition (OCR) systems. This research addresses this challenge of data scarcity by leveraging parallel textline images - pairs of original Kindai text and their counterparts in contemporary Japanese fonts - to augment training datasets. We introduce a distance-based objective function that minimizes the gap between self-attention features of the parallel image pairs. Specifically, we explore Euclidean distance and Maximum Mean Discrepancy (MMD) as domain adaptation metrics. Experimental results demonstrate that our method reduces the character error rate (CER) by 2.23% and 3.94% over a Transformer-based OCR baseline when using Euclidean distance and MMD, respectively. Furthermore, our approach improves the discriminative quality of self-attention representations, leading to more effective OCR performance for historical documents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。