无监督模型学会识别汉语四声,模拟人类学语过程。
Unsupervised Learning and Representation of Mandarin Tonal Categories by a Generative CNN
- 用生成式CNN在无标签数据下学习汉语声调分类
- 男性语音训练的模型稳定编码声调差异,且符合人类习得阶段
- 通过卷积层追踪声调表征,提升模型可解释性
本文提出一种无监督建模方法,用于模拟人类语言习得中的声调学习。声调是语言中计算最复杂的习得目标之一。我们论证,一个真实的生成模型(ciwGAN)可在无任何标注数据的情况下,将类别变量与普通话声调类别关联起来。所有三个训练模型在基频(F0)上均表现出统计显著的差异。仅使用男性语音训练的模型始终稳定编码声调。结果表明,该模型不仅学会了普通话的声调对立,还习得了一个与人类语言学习者发展阶段相符的系统。此外,本文还提出一种追踪内部卷积层中声调表征的方法,表明语言学工具可助力深度学习模型的可解释性,并可用于神经实验。
原文摘要 · Abstract (English)
This paper outlines the methodology for modeling tonal learning in fully unsupervised models of human language acquisition. Tonal patterns are among the computationally most complex learning objectives in language. We argue that a realistic generative model of human language (ciwGAN) can learn to associate its categorical variables with Mandarin Chinese tonal categories without any labeled data. All three trained models showed statistically significant differences in F0 across categorical variables. The model trained solely on male tokens consistently encoded tone. Our results sug- gest that not only does the model learn Mandarin tonal contrasts, but it learns a system that corresponds to a stage of acquisition in human language learners. We also outline methodology for tracing tonal representations in internal convolutional layers, which shows that linguistic tools can contribute to interpretability of deep learning and can ultimately be used in neural experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。