arXiv:2608.00096cs.CVcs.AI2026-08

用语义对比学习提升汉字图像识别,尤其改善罕见字表现

Logographic Character Visual Pretraining via Semantic-based Contrastive Learning

论文配图:Logographic Character Visual Pretraining via Semantic-based Contrastive Learning
图 1 · 摘自论文原文
  • 结合视觉与上下文语义,设计新型对比预训练策略
  • 在不平衡数据下识别准确率显著优于现有方法
  • 适合处理中文等表意文字的图像识别任务

当前基于深度学习的字符视觉研究(如文本识别、字符图像去噪、古籍文本补全)在学习、管理与利用字符资源方面提供了新方案。然而,这些方法的表现仅在大规模且平衡的数据集上达到峰值,而真实世界中的字符数据集,尤其是汉字等表意文字,往往存在数据分布不均的问题,这主要源于字符使用频率差异以及新字不断出现。本文提出一种新的汉字识别方法,引入多模态学习框架,融合字符的视觉语义与上下文语义。设计了一种新颖的预训练策略,通过语言模型提取每个字符的上下文语义,以增强深层视觉表征,尤其适用于存在类别不平衡和稀有实例的数据集。我们在多个数据集上进行实验,评估所提方法,并通过多个下游任务验证对比预训练策略的有效性。实验结果表明,该方法在性能上优于现有最先进方法。

原文摘要 · Abstract (English)

Current deep learning-based character vision studies, e.g., text recognition, character image denoising, and historical text completion, are offering new solutions for learning, managing, and utilizing character resources. However, the performance of these studies peaks only with large and balanced datasets, which is a rarity with real-world character datasets, especially for logographic character languages, e.g., Chinese. The imbalance in data distribution of logographic characters is a common issue due to differences in character usage frequency and new characters being continuously created. In this paper, we propose a novel method for logographic character recognition, which introduces a multi-modal learning approach using visual semantics and contextual semantics of characters. A novel pre-training strategy is designed to enhance deep visual representations, especially for datasets suffering from issues of imbalanced and rare instances, by extracting the contextual semantics of each character from the corresponding language models. We conduct experiments across various datasets to evaluate our character recognition method and further validate the contrastive pre-training strategy by several downstream tasks. Experimental results demonstrate the superiority of our method compared to state-of-the-art methods.

字符识别视觉预训练表意文字

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。