arXiv:2506.03798cs.CV2025-06被引 2

提出无监督分解汉字的潜空间模型,实现零样本识别。

CoLa: Chinese Character Decomposition with Compositional Latent Components

  • 用潜变量模型自动学习汉字组成成分,不依赖人工拆分。
  • 在零样本场景下准确率超越现有方法,历史文献也能泛化。
  • 适合研究汉字认知、跨数据集识别与可解释模型的学者。

人类能将汉字分解为构成部件并重新组合以识别生字,体现组合性与元学习两大认知原则,提供高效泛化的归纳偏置。这对解决汉字数据集长尾分布导致的零样本识别问题至关重要。现有方法虽通过预定义部首或笔画建模组合性,但忽略元学习能力,泛化受限。受此启发,我们提出一种深度潜变量模型 CoLa,无需人工设定分解方案,直接学习汉字的组合潜成分。通过潜空间中成分对比实现识别与匹配,支持零样本字符识别。实验表明,CoLa 在部首零样本汉字识别任务上优于以往方法;可视化显示所学成分具有可解释性;且在历史文献上训练后仍能分析甲骨文成分,展现强大跨数据集泛化能力。

原文摘要 · Abstract (English)

Humans can decompose Chinese characters into compositional components and recombine them to recognize unseen characters. This reflects two cognitive principles: Compositionality, the idea that complex concepts are built on simpler parts; and Learning-to-learn, the ability to learn strategies for decomposing and recombining components to form new concepts. These principles provide inductive biases that support efficient generalization. They are critical to Chinese character recognition (CCR) in solving the zero-shot problem, which results from the common long-tail distribution of Chinese character datasets. Existing methods have made substantial progress in modeling compositionality via predefined radical or stroke decomposition. However, they often ignore the learning-to-learn capability, limiting their ability to generalize beyond human-defined schemes. Inspired by these principles, we propose a deep latent variable model that learns Compositional Latent components of Chinese characters (CoLa) without relying on human-defined decomposition schemes. Recognition and matching can be performed by comparing compositional latent components in the latent space, enabling zero-shot character recognition. The experiments illustrate that CoLa outperforms previous methods in both character the radical zero-shot CCR. Visualization indicates that the learned components can reflect the structure of characters in an interpretable way. Moreover, despite being trained on historical documents, CoLa can analyze components of oracle bone characters, highlighting its cross-dataset generalization ability.

汉字识别潜变量模型零样本学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。