首个针对古尼泊尔手稿的端到端文字识别系统,准确率达4.9%
Digitizing Nepal's Written Heritage: A Comprehensive HTR Pipeline for Old Nepali Manuscripts
- 采用编码器-解码器架构,结合数据增强提升识别效果
- 最佳模型字符错误率降至4.9%,显著优于现有水平
- 开源训练代码与配置,助力低资源历史文字研究
本文首次提出针对古尼泊尔语的手写文字识别(HTR)全流程解决方案。该语言历史悠久但资源匮乏,研究采用行级转录方式,系统探索了编码器-解码器结构与数据驱动技术以提升识别准确率。最佳模型在未公开评估集上实现4.9%的字符错误率(CER)。研究还实现了多种解码策略并分析了词元级混淆情况,以深入理解模型行为与错误模式。尽管评估数据集保密,论文已公开训练代码、模型配置及评估脚本,旨在推动对低资源历史文字的HTR研究。
原文摘要 · Abstract (English)
This paper presents the first end-to-end pipeline for Handwritten Text Recognition (HTR) for Old Nepali, a historically significant but low-resource language. We adopt a line-level transcription approach and systematically explore encoder-decoder architectures and data-centric techniques to improve recognition accuracy. Our best model achieves a Character Error Rate (CER) of 4.9\%. In addition, we implement and evaluate decoding strategies and analyze token-level confusions to better understand model behavior and error patterns. Although the evaluation dataset is confidential, we release our training code, model configurations, and evaluation scripts to support further research on HTR for low-resource historical scripts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。