统一识别13种印度手写文字,支持古籍数字化与自动编目。
UniLipi: A Unified Multi-Script OCR for Historical Indic Manuscripts

- 单一模型联合训练13种印度手写体,无需逐个定制。
- 在低资源下仍保持高精度,依赖合成数据减少真实标注需求。
- 可推广至现代手写及非印度文字,适合文化遗产数字化项目。
手写印度古籍的光学字符识别对大规模数字化和计算访问手稿遗产至关重要。然而,现有方法通常针对单一文字体系开发,需大量脚本特异性调优,限制了跨多样化藏品的可扩展性和实际部署。我们提出UniLipi,一种在单一框架内联合训练13种印度手写文字的统一多文字OCR模型。该模型直接处理真实古籍中的复杂情况,包括极端线段几何变化、长宽差异大以及因孔洞、污渍或插图导致的部分中断。为在超低资源条件下有效运行,模型采用脚本感知的合成手写古籍数据生成,大幅减少对大规模真实标注数据的依赖。除转录外,UniLipi还能预测文字体系身份和每行原生字符数,支持实际古籍编目工作流程。此外,我们证明UniLipi可作为有效的基础预训练模型,其学习表征在当代印度手写体上表现良好,并拓展至藏文、意大利文、拉丁文和中文等非印度文字,具有广泛适用性。
原文摘要 · Abstract (English)
Optical character recognition (OCR) for handwritten Indic manuscripts is essential for large-scale digitization and computational access to manuscript heritage. However, existing approaches are typically developed for one script at a time and require substantial script-specific customization. This limits scalability and practical deployment across diverse collections. We present UniLipi, a unified multi-script OCR model for handwritten Indic manuscripts trained jointly across 13 Indic scripts within a single framework. UniLipi directly handles realistic manuscript conditions, including extreme variation in line geometry, large variation in line length, and partial interruptions caused by non-textual manuscript entities such as holes, stains, or pictorial illustrations. To operate effectively under ultra low-resource conditions, the model leverages script-aware synthetic manuscript data generation, substantially reducing reliance on large volumes of real annotated data. Beyond historical manuscripts, we show that UniLipi serves as an effective foundational pretrained model. Specifically, its learned representations enable good OCR performance for contemporary Indic handwriting and extend to several non-Indic scripts, including Tibetan, Italian, Latin, and Chinese scripts. In addition to transcription, UniLipi predicts script identity and per-line native character counts, supporting practical manuscript cataloging workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。