自监督模型通过声学不变性实现语音模仿,模拟婴儿学习发音的机制。
From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation model
- 用wav2vec 2.0中间层特征替代传统MFCC,提升发音参数学习效果
- 学习到的发音轨迹与人类模式高度相关,能区分发音部位
- 揭示感知学习引导发音发展的计算机制,适合语言认知研究者
人类婴儿在没有明确指导的情况下,需将极富变化的听觉输入映射为合适的发音动作。本文提出一个自监督语音模仿模型,包含特征提取器、逆向映射模型和语音合成器。实验表明,预训练的wav2vec 2.0模型中间层特征显著优于传统MFCC,能有效支持发音参数学习。该模型生成的发音轨迹与人类行为高度一致,可区分不同发音部位,并产生可懂语音。成功的关键在于特征同时具备语音可辨性和说话人不变性——这正是自监督表示学习的核心特性。研究结果为‘感知学习引导发音发展’的发育理论提供了计算支持,揭示了婴儿如何克服复杂映射难题习得语音产出能力。
原文摘要 · Abstract (English)
Human infants face a formidable challenge in speech acquisition: mapping extremely variable acoustic inputs into appropriate articulatory movements without explicit instruction. We present a computational model that addresses the acoustic-to-articulatory mapping problem through self-supervised learning. Our model comprises a feature extractor that transforms speech into latent representations, an inverse model that maps these representations to articulatory parameters, and a synthesizer that generates speech outputs. Experiments conducted in both single- and multi-speaker settings reveal that intermediate layers of a pre-trained wav2vec 2.0 model provide optimal representations for articulatory learning, significantly outperforming MFCC features. These representations enable our model to learn articulatory trajectories that correlate with human patterns, discriminate between places of articulation, and produce intelligible speech. Critical to successful articulatory learning are representations that balance phonetic discriminability with speaker invariance -- precisely the characteristics of self-supervised representation learning models. Our findings provide computational evidence consistent with developmental theories proposing that perceptual learning of phonetic categories guides articulatory development, offering insights into how infants might acquire speech production capabilities despite the complex mapping problem they face.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。