arXiv:2510.04003cs.CVcs.CL2025-10被引 4

微调PaddleOCRv5提升古越南汉字识别准确率

Enhancing OCR for Sino-Vietnamese Language Processing via Fine-tuned PaddleOCRv5

  • 用古越南汉文手稿微调PaddleOCRv5文本识别模块
  • 准确率从37.5%提升至50.0%,尤其在噪声图像下表现更好
  • 适合历史文献数字化与中越语义对齐研究者使用

识别与处理古典中文(汉喃)文本对于数字化越南历史文献及开展跨语言语义研究至关重要。然而,现有OCR系统在处理老旧扫描件、非标准字形和手写体变化时表现不佳。本文提出一种针对PaddleOCRv5的微调方法,通过使用精选的古代越南汉语手稿子集重新训练文本识别模块,并构建完整的训练流程,涵盖预处理、LMDB转换、评估与可视化。实验结果显示,相较于基础模型,准确率从37.5%显著提升至50.0%,尤其在噪声图像条件下效果更优。此外,我们开发了一个交互式演示系统,可直观对比微调前后识别结果,助力汉喃语义对齐、机器翻译与历史语言学研究。演示地址:https://huggingface.co/spaces/MinhDS/Fine-tuned-PaddleOCRv5

原文摘要 · Abstract (English)

Recognizing and processing Classical Chinese (Han-Nom) texts play a vital role in digitizing Vietnamese historical documents and enabling cross-lingual semantic research. However, existing OCR systems struggle with degraded scans, non-standard glyphs, and handwriting variations common in ancient sources. In this work, we propose a fine-tuning approach for PaddleOCRv5 to improve character recognition on Han-Nom texts. We retrain the text recognition module using a curated subset of ancient Vietnamese Chinese manuscripts, supported by a full training pipeline covering preprocessing, LMDB conversion, evaluation, and visualization. Experimental results show a significant improvement over the base model, with exact accuracy increasing from 37.5 percent to 50.0 percent, particularly under noisy image conditions. Furthermore, we develop an interactive demo that visually compares pre- and post-fine-tuning recognition results, facilitating downstream applications such as Han-Vietnamese semantic alignment, machine translation, and historical linguistics research. The demo is available at https://huggingface.co/spaces/MinhDS/Fine-tuned-PaddleOCRv5

OCR古文献汉喃微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。