arXiv:2501.01981cs.CVeess.IV2025-01被引 8

用深度学习识别古代婆罗米文字,移动端模型准确率达95.94%。

Optical Character Recognition using Convolutional Neural Networks for Ashokan Brahmi Inscriptions

  • 基于预训练CNN模型,结合迁移学习与图像增强提升识别能力。
  • MobileNet在验证集上达到95.94%准确率,损失仅0.129。
  • 适合古文字数字化、考古学与文化遗产保护领域研究者参考。

本研究构建了用于识别阿育王婆罗米文字的光学字符识别(OCR)系统,采用卷积神经网络(CNN)进行建模。通过大规模字符图像数据集训练模型,并运用数据增强和图像预处理技术降低噪声,实现文本行与字符的精准分割。重点对比了LeNet、VGG-16与MobileNet三个预训练CNN模型,均采用迁移学习策略适配婆罗米文字数据。结果表明,MobileNet表现最优,验证准确率达95.94%,验证损失为0.129。研究深入分析了MobileNet的实现过程,并探讨了其在铭文学领域的应用价值。该方法对古代文字的数字化保存具有重要意义。

原文摘要 · Abstract (English)

This research paper delves into the development of an Optical Character Recognition (OCR) system for the recognition of Ashokan Brahmi characters using Convolutional Neural Networks. It utilizes a comprehensive dataset of character images to train the models, along with data augmentation techniques to optimize the training process. Furthermore, the paper incorporates image preprocessing to remove noise, as well as image segmentation to facilitate line and character segmentation. The study mainly focuses on three pre-trained CNNs, namely LeNet, VGG-16, and MobileNet and compares their accuracy. Transfer learning was employed to adapt the pre-trained models to the Ashokan Brahmi character dataset. The findings reveal that MobileNet outperforms the other two models in terms of accuracy, achieving a validation accuracy of 95.94% and validation loss of 0.129. The paper provides an in-depth analysis of the implementation process using MobileNet and discusses the implications of the findings. The use of OCR for character recognition is of significant importance in the field of epigraphy, specifically for the preservation and digitization of ancient scripts. The results of this research paper demonstrate the effectiveness of using pre-trained CNNs for the recognition of Ashokan Brahmi characters.

OCR古文字深度学习迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。