用深度学习提升中世纪手写拉丁文文献的文本识别准确率。
Application of deep learning approaches for medieval historical documents transcription
- 针对中世纪手写体设计专用深度学习流程,结合文档特性优化识别。
- 在9至11世纪拉丁文手稿上实现高精度文本提取,关键指标均超基准。
- 开源代码与数据集,适合历史文献数字化研究者使用。
手写文本识别与光学字符识别在现代文本处理中表现优异,但在中世纪拉丁文手稿上效率显著下降。本文提出一种深度学习方法,用于从公元9至11世纪的拉丁文手写文献中提取文本信息。该方法充分考虑了中世纪文献的固有特征。论文简要介绍了历史文献转录领域,对原始数据进行了初步分析,并综述了相关研究。详细描述了数据集构建流程,用于后续模型训练。同时提供了处理后数据的解释性分析。论文阐述了从文档图像中提取文本的完整深度学习流水线:从目标检测到基于分类模型和词图像嵌入的单词识别。实验报告了召回率、精确率、F1分数、交并比、混淆矩阵及平均字符串距离等指标,并附有相应图表。实现代码已发布于GitHub仓库。
原文摘要 · Abstract (English)
Handwritten text recognition and optical character recognition solutions show excellent results with processing data of modern era, but efficiency drops with Latin documents of medieval times. This paper presents a deep learning method to extract text information from handwritten Latin-language documents of the 9th to 11th centuries. The approach takes into account the properties inherent in medieval documents. The paper provides a brief introduction to the field of historical document transcription, a first-sight analysis of the raw data, and the related works and studies. The paper presents the steps of dataset development for further training of the models. The explanatory data analysis of the processed data is provided as well. The paper explains the pipeline of deep learning models to extract text information from the document images, from detecting objects to word recognition using classification models and embedding word images. The paper reports the following results: recall, precision, F1 score, intersection over union, confusion matrix, and mean string distance. The plots of the metrics are also included. The implementation is published on the GitHub repository.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。