arXiv:2506.10855cs.CLeess.AS2025-06被引 3

研究不同语言下语音模型如何编码音素、声调和说话人信息

Analyzing the relationships between pretraining language, phonetic, tonal, and speaker information in self-supervised speech models

  • 用探针分类器和几何分析法,研究多语言wav2vec2模型的表征结构
  • 音素、声调和说话人信息在各层中基本正交,且表现模式相似
  • 跨语言训练时,音素与声调有轻微匹配优势,但说话人信息无此现象

自监督语音模型的分析已开始揭示其如何表示不同类型的信息,但几乎都集中在英语。本文研究了在四种不同语言上预训练的wav2vec2模型,对匹配与非匹配语言语音的表征情况。通过探针分类器和几何分析方法,考察音素、词汇声调和说话人信息的编码方式。结果表明,对于所有预训练和测试语言,音素、声调和说话人信息的子空间基本正交;逐层探针准确率模式相似,后期层中匹配语言的音素与声调探针有小幅优势(但说话人探针无),说明wav2vec2学习到的表征结构在很大程度上独立于预训练语音材料。

原文摘要 · Abstract (English)

Analyses of self-supervised speech models have begun to reveal where and how they represent different types of information. However, almost all analyses have focused on English. Here, we examine how wav2vec2 models trained on four different languages encode both language-matched and non-matched speech. We use probing classifiers and geometric analyses to examine how phones, lexical tones, and speaker information are represented. We show that for all pretraining and test languages, the subspaces encoding phones, tones, and speakers are largely orthogonal, and that layerwise patterns of probing accuracy are similar, with a relatively small advantage for matched-language phone and tone (but not speaker) probes in the later layers. Our findings suggest that the structure of representations learned by wav2vec2 is largely independent of the speech material used during pretraining.

自监督语音表征分析多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。