arXiv:2508.10356cs.CVcs.CL2025-08被引 2

用深度学习提升多语言历史文本的识别准确率

Improving OCR for Historical Texts of Multiple Languages

  • 融合数据增强与先进模型提升古希伯来文识别
  • 结合语义分割与置信度伪标签,显著优化会议记录识别
  • 采用残差网络与CTC损失,有效捕捉手写英文时序特征

本文基于深度学习技术,在三个任务中推进了光学字符识别(OCR)与文档版面分析。针对死海古卷中的古代希伯来文片段,通过大量数据增强并使用Kraken和TrOCR模型提升字符识别效果;在16至18世纪会议决议文件识别任务中,采用集成DeepLabV3+语义分割与双向LSTM的卷积循环神经网络(CRNN),结合置信度伪标签优化模型性能;对于现代英文手写体识别,使用带ResNet34编码器的CRNN,并以连接时序分类(CTC)损失函数训练,有效捕捉序列依赖关系。研究为多语言历史文本处理提供了可行方法与未来方向。

原文摘要 · Abstract (English)

This paper presents our methodology and findings from three tasks across Optical Character Recognition (OCR) and Document Layout Analysis using advanced deep learning techniques. First, for the historical Hebrew fragments of the Dead Sea Scrolls, we enhanced our dataset through extensive data augmentation and employed the Kraken and TrOCR models to improve character recognition. In our analysis of 16th to 18th-century meeting resolutions task, we utilized a Convolutional Recurrent Neural Network (CRNN) that integrated DeepLabV3+ for semantic segmentation with a Bidirectional LSTM, incorporating confidence-based pseudolabeling to refine our model. Finally, for modern English handwriting recognition task, we applied a CRNN with a ResNet34 encoder, trained using the Connectionist Temporal Classification (CTC) loss function to effectively capture sequential dependencies. This report offers valuable insights and suggests potential directions for future research.

OCR历史文本深度学习手写识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。