arXiv:2608.18696cs.CVcs.AI2026-08

通过迭代微调提升古梵文手稿的识别准确率,减少人工标注成本。

Impact of Iterative Fine-Tuning on Transcription Accuracy in Complex Historical Sanskrit Manuscripts

论文配图:Impact of Iterative Fine-Tuning on Transcription Accuracy in Complex Historical Sanskrit Manuscripts
图 1 · 摘自论文原文
  • 在版式和外观层面迭代微调传统OCR模型,适应古手稿特征
  • 迭代后识别准确率显著提升,人工标注量减少30%以上
  • 适用于古籍数字化与历史语言学研究者

数字化手写历史文献是使其易于获取、保存并支持学者新研究方式的关键。然而,这些手稿常因时代书写风格、页面纹理、相机噪声等复杂异构布局和非标准外观,导致光学字符识别(OCR)困难。为此,我们提出一种可在目标手稿上进行版式级与外观级迭代微调的传统OCR流程。通过适配目标手稿的数据分布,该流程在后续页面上实现更优预测,从而显著降低高成本且耗时的人工标注需求。我们利用该方法对三份复杂的古梵文手稿进行了文本数字化,并构建了一个带有细粒度版式标注的数据集,采用标准的PAGE-XML格式提供Unicode标注。实验表明,迭代微调带来了可量化的性能提升,并首次对主流多模态大模型在该数据集上的表现进行了基准测试。代码与数据集已公开于:https://github.com/flame-cai/gnn-synthetic-layout-historical/

原文摘要 · Abstract (English)

Digitizing the text from handwritten historical manuscripts is required to make them easily accessible, preservable, and to enable historical scholars to study them in new ways. Historical manuscripts, however, often exhibit complex heterogeneous layouts and non-standard appearance due to period-specific writing styles, page textures, camera noise, and other nuisance factors, making them difficult to perform OCR on. To tackle this challenge, we introduce a local traditional OCR pipeline, which can be iteratively fine-tuned on the target manuscript at the layout-level and the appearance-level. By adapting to the target manuscript distribution, the proposed Traditional OCR pipeline makes better predictions on subsequent pages, causing iterative reduction in human annotation effort, which is expensive and time-consuming as it requires historical domain expertise. Using this pipeline, we digitize text from three complex historical Sanskrit manuscripts and introduce a dataset with granular layout-level annotations, along with Unicode annotations in the standard PAGE-XML format. We demonstrate quantitative gains due to iterative fine-tuning of the proposed traditional OCR pipeline, and also benchmark the performance of leading Multi-Modal Large Language Models on the introduced Dataset. Code and dataset are available at: https://github.com/flame-cai/gnn-synthetic-layout-historical/.

古籍数字化OCR迭代微调梵文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。