arXiv:2602.03913cs.CVcs.AI2026-02

通过信息熵动态聚焦关键部件,提升零样本手写汉字识别准确率

Entropy-Aware Structural Alignment for Zero-Shot Handwritten Chinese Character Recognition

  • 用信息熵调节位置嵌入,自动突出重要字根
  • 构建双视角字根树提取多粒度结构特征,准确率达55.04%
  • 仅需每类1个样本即可达到92.41%准确率,适合小样本场景

零样本手写汉字识别旨在利用部件语义组合识别未见字符。现有方法常将字符视为扁平部件序列,忽略层级结构与组件信息密度差异。为此,提出熵感知结构对齐网络,通过信息论建模弥合视觉-语义鸿沟。首先引入信息熵先验,通过乘法交互动态调制位置嵌入,充当显著性检测器,优先关注判别性部件而非常见成分。其次构建双视角字根树,提取多粒度结构特征,并通过自适应Sigmoid门控网络融合,编码全局布局与局部空间角色。最后设计Top-K语义特征融合机制,利用语义邻居质心增强解码,通过特征级共识有效纠正视觉歧义。大量实验表明,该方法在ICDAR 2013数据集上(m=1500)达到55.04%准确率,显著优于现有基于CLIP的基线。此外,框架具备极强数据效率,仅需每类一个支持样本即达92.41%准确率。

原文摘要 · Abstract (English)

Zero-shot Handwritten Chinese Character Recognition (HCCR) aims to recognize unseen characters by leveraging radical-based semantic compositions. However, existing approaches often treat characters as flat radical sequences, neglecting the hierarchical topology and the uneven information density of different components. To address these limitations, we propose an Entropy-Aware Structural Alignment Network that bridges the visual-semantic gap through information-theoretic modeling. First, we introduce an Information Entropy Prior to dynamically modulate positional embeddings via multiplicative interaction, acting as a saliency detector that prioritizes discriminative roots over ubiquitous components. Second, we construct a Dual-View Radical Tree to extract multi-granularity structural features, which are integrated via an adaptive Sigmoid-based gating network to encode both global layout and local spatial roles. Finally, a Top-K Semantic Feature Fusion mechanism is devised to augment the decoding process by utilizing the centroid of semantic neighbors, effectively rectifying visual ambiguities through feature-level consensus. Extensive experiments demonstrate that our method establishes new state-of-the-art performance, achieving an accuracy of 55.04\% on the ICDAR 2013 dataset ($m=1500$), significantly outperforming existing CLIP-based baselines in the challenging zero-shot setting. Furthermore, the framework exhibits exceptional data efficiency, demonstrating rapid adaptability with minimal support samples, achieving 92.41\% accuracy with only one support sample per class.

零样本识别结构建模小样本学习汉字识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。