arXiv:2506.06761cs.LGcs.CV2025-06被引 3

用模型编辑提升对低资源文字的识别能力,无需大量数据重训练。

The OCR Quest for Generalization: Learning to recognize low-resource alphabets with model editing

  • 通过模型编辑技术快速融入未见字符集,替代传统微调。
  • 在历史密文和非拉丁字母上实现显著性能提升。
  • 适合需要快速适配新语言或古文字的文档识别场景。

在多样化领域中实现鲁棒的识别系统对实际应用至关重要。尽管通常假定数据充足,但低资源语言(如古文献、非西方语言)因代表性不足,常被大型预训练和基础方法排除在外。本文旨在构建能快速泛化至新数据分布(如新字母表)的模型,超越集中式微调策略。我们利用模型编辑的最新进展,增强对未见脚本(低资源学习)的适应能力。与最先进的元学习方法相比,我们在稀疏数据分布下展示了领域融合的有效性,且不依赖整体分布关系或原型生成。即使使用完全相同的训练数据,实验表明在迁移学习至新字母表及面对挑战性领域偏移(包括历史密文和非拉丁文字)的域外评估中均有显著性能提升。本研究提出一种新方法,使模型能轻松接纳低代表性的字母表,从而拓展文档识别的应用范围至更广泛的文化与语境。

原文摘要 · Abstract (English)

Achieving robustness in recognition systems across diverse domains is crucial for their practical utility. While ample data availability is usually assumed, low-resource languages, such as ancient manuscripts and non-western languages, tend to be kept out of the equations of massive pretraining and foundational techniques due to an under representation. In this work, we aim for building models which can generalize to new distributions of data, such as alphabets, faster than centralized fine-tune strategies. For doing so, we take advantage of the recent advancements in model editing to enhance the incorporation of unseen scripts (low-resource learning). In contrast to state-of-the-art meta-learning, we showcase the effectiveness of domain merging in sparse distributions of data, with agnosticity of its relation to the overall distribution or any other prototyping necessity. Even when using the same exact training data, our experiments showcase significant performance boosts in \textbf{transfer learning} to new alphabets and \textbf{out-of-domain evaluation} in challenging domain shifts, including historical ciphered texts and non-Latin scripts. This research contributes a novel approach into building models that can easily adopt under-represented alphabets and, therefore, enable document recognition to a wider set of contexts and cultures.

OCR模型编辑低资源文字识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。