语音编码器的学习目标决定非母语发音辨识能力的下降
The Learning Objective Governs Perceptual Narrowing: A Cross-Lingual, Layer-Wise, Ten-Seed Study of Self-Supervised Speech Encoders

- 用不同学习目标训练模型,发现重建目标削弱非母语辨识,预测目标增强它
- 重建目标使第一层发音辨识低于输入,预测目标则高于输入,差异达0.051
- 需十次种子实验才能稳定观察到效果,三次种子易误判结果
感知窄化——婴儿出生后第一年对非母语音素辨识能力逐渐丧失——是经典发展现象,但其背后的驱动机制仍不明确。我们使用约700万参数的Transformer编码器,在儿童导向语音和朗读语音上进行训练,并在英语、法语和普通话中评估音素ABX任务,共进行十次独立种子实验。结果表明:(1) 学习目标决定跨语言迁移方向:重建(掩码梅尔频谱预测)导致非母语辨识下降,预测(帧对比)则提升它,第一层普通话ABX差异为+0.051(p=3×10⁻⁸),20次运行一致;(2) 辨识下降由内部难度梯度主导,语言匹配效应较小(匹配与不匹配间差+0.022,p=10⁻⁴,四层均显著);(3) 相较于语言对称的原始梅尔频谱基线,重建使第一层辨识能力低于输入,预测则高于输入;(4) 朗读语音导致的非母语衰退速度是儿童导向语音的3.6倍;(5) 常规三种子实验预算无法可靠检测此效应,十种子下显著的结果仅在70%的三种子子集中被识别;(6) 六种目标配置(锐化、压缩、整合及其组合,以及两种词级语义锚定)均未能产生完整的发育特征(母语提升且非母语下降):单一目标作用于共享表示,使两语言同步变化。结论:学习目标而非架构,是塑造窄化式表征变化的首要决定因素。
原文摘要 · Abstract (English)
Perceptual narrowing---the developmental loss of non-native phoneme discrimination in the first year of life \citep{werker1984}---is a canonical developmental finding, yet \emph{what learning objective produces it} remains open. We train a \(\sim\)7\,M-parameter Transformer encoder on child-directed and read speech and evaluate phoneme ABX in English, French, and Mandarin over ten seeds, the seed as the unit of replication. Six results. \textbf{(1)}~The objective sets the direction of cross-lingual transfer: reconstruction (masked mel-prediction) degrades non-native discrimination, prediction (frame-contrastive) improves it---a same-encoder, same-data gap of \(+0.051\) in first-layer Mandarin ABX (\(p=3\times10^{-8}\)), unanimous in sign across twenty runs. \textbf{(2)}~That decline combines a large arm-intrinsic difficulty gradient with a smaller language-specialization effect (matched vs.\ mismatched \(+0.022\), \(p=10^{-4}\), all four layers). \textbf{(3)}~Against a language-symmetric raw-mel floor, reconstruction pushes the first layer \emph{below} the discriminability of its input; prediction pushes it \emph{above}. \textbf{(4)}~Read speech gives a \(3.6\times\) steeper non-native decline than child-directed speech. \textbf{(5)}~The customary three-seed budget cannot see this reliably: an effect unambiguous at ten seeds is called significant by as few as 70\% of three-seed subsets. \textbf{(6)}~Six objective configurations---sharpening, compression, consolidation, their composition, and word-level semantic grounding in two forms---fail to produce the full developmental signature (native improves \emph{and} non-native declines): a single objective moves both languages the same way because it acts on a shared representation. We conclude that the objective, not the architecture, is the first-order determinant of narrowing-shaped representational change.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。