提出首个通用高棉文文本识别框架,支持印刷、手写、场景多种文本形态。
Towards Universal Khmer Text Recognition
- 设计模态感知自适应特征选择机制,动态适配不同文本类型。
- 在多模态数据上实现最优性能,尤其提升低资源手写文本识别率。
- 开源首个高棉文通用文本识别基准,助力后续研究发展。
高棉语是一种低资源语言,其复杂文字系统给光学字符识别(OCR)带来巨大挑战。尽管印刷体文档识别因可用数据集而取得进展,但手写体和场景文本的识别仍受限于数据稀缺。为每种模态训练专用模型无法实现跨模态迁移学习,且导致内存开销大、输入路由易出错。简单合并多模态数据进行训练,因分布不均常导致低频模态性能下降。为此,本文提出通用高棉文文本识别(UKTR)框架,可处理多种文本模态。核心是提出一种新型模态感知自适应特征选择(MAFS)技术,根据输入图像模态动态调整视觉特征,增强跨模态识别鲁棒性。大量实验表明,该模型达到当前最优性能。此外,我们构建了首个全面的通用高棉文文本识别基准,并向社区开放,以推动后续研究。相关数据集与模型可通过受控仓库获取。
原文摘要 · Abstract (English)
Khmer is a low-resource language characterized by a complex script, presenting significant challenges for optical character recognition (OCR). While document printed text recognition has advanced because of available datasets, performance on other modalities, such as handwritten and scene text, remains limited by data scarcity. Training modality-specific models for each modality does not allow cross-modality transfer learning, from which modalities with limited data could otherwise benefit. Moreover, deploying many modality-specific models results in significant memory overhead and requires error-prone routing each input image to the appropriate model. On the other hand, simply training on a combined dataset with a non-uniform data distribution across different modalities often leads to degraded performance on underrepresented modalities. To address these, we propose a universal Khmer text recognition (UKTR) framework capable of handling diverse text modalities. Central to our method is a novel modality-aware adaptive feature selection (MAFS) technique designed to adapt visual features according to a particular input image modality and enhance recognition robustness across modalities. Extensive experiments demonstrate that our model achieves state-of-the-art (SoTA) performance. Furthermore, we introduce the first comprehensive benchmark for universal Khmer text recognition, which we release to the community to facilitate future research. Our datasets and models can be accessible via this gated repository\footnote{in review}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。