用对比学习提升濒危语言文字识别准确率,仅需少量标注数据即可实现高精度识别。
VOLTAGE: A Versatile Contrastive Learning based OCR Methodology for ultra low-resource scripts through Auto Glyph Feature Extraction
- 基于对比学习与自动生成字形特征,实现低资源文字的自动标注。
- 对印刷体和手写体塔克里文识别准确率达95%和87%。
- 适用于多种印地语系文字,尤其适合濒危语言数字化保护。
联合国教科文组织将全球7000种语言中的2500种列为濒危语言。语言消亡意味着传统智慧、民间文学及社群精髓的丧失,亟需推动其数字包容以避免灭绝。低资源语言面临更高灭绝风险,而缺乏无监督光学字符识别(OCR)方法是阻碍其数字化的关键因素。本文提出VOLTAGE——一种基于对比学习的OCR方法,通过自动生成字形特征实现聚类标注,并利用图像变换与生成对抗网络扩充标注数据以增强多样性与数量。该方法以16至20世纪印度喜马拉雅地区使用的塔克里文(Takri)为设计基础,同时在其他印地语系文字(包括高低资源语种)上验证其普适性。实验表明,对印刷体和手写体塔克里文的识别准确率分别达到95%和87%。我们还进行了基线与消融研究,并构建下游应用场景,证明该方法在实际中的有效性。
原文摘要 · Abstract (English)
UNESCO has classified 2500 out of 7000 languages spoken worldwide as endangered. Attrition of a language leads to loss of traditional wisdom, folk literature, and the essence of the community that uses it. It is therefore imperative to bring digital inclusion to these languages and avoid its extinction. Low resource languages are at a greater risk of extinction. Lack of unsupervised Optical Character Recognition(OCR) methodologies for low resource languages is one of the reasons impeding their digital inclusion. We propose VOLTAGE - a contrastive learning based OCR methodology, leveraging auto-glyph feature recommendation for cluster-based labelling. We augment the labelled data for diversity and volume using image transformations and Generative Adversarial Networks. Voltage has been designed using Takri - a family of scripts used in 16th to 20th century in the Himalayan regions of India. We present results for Takri along with other Indic scripts (both low and high resource) to substantiate the universal behavior of the methodology. An accuracy of 95% for machine printed and 87% for handwritten samples on Takri script has been achieved. We conduct baseline and ablation studies along with building downstream use cases for Takri, demonstrating the usefulness of our work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。