arXiv:2606.31527eess.AScs.SD2026-06

用双语发音数据验证自监督语音模型的跨语言发音表征能力

How Bilingual Are SSL Speech Models? Cross-Lingual Probing of Articulatory Encoding with Finnish and Russian EMA

论文配图:How Bilingual Are SSL Speech Models? Cross-Lingual Probing of Articulatory Encoding with Finnish and Russian EMA
图 1 · 摘自论文原文
  • 基于双语者发音动作数据,分析自监督模型对发音运动的编码能力
  • 仅用5分钟训练数据,预测相关性最高达0.68,中间层表现最佳
  • 模型对舌部运动预测更准,且在不同语言熟练度和口语类型间泛化良好

自监督语音模型能从原始音频中捕捉丰富的语音、语调和声学模式,但其在不同语言间如何编码发音信息尚不明确。我们利用双语芬兰语-俄语使用者的发音动作追踪(EMA)数据,评估了自监督模型隐状态表示与发音运动之间的跨语言相关性。模型在仅约5分钟训练数据下即达到较强预测性能(皮尔逊相关系数最高达0.68),多语言模型优于单语言模型。中间层对发音特征的编码最为有效,且舌部运动比唇部运动更易预测。我们还考察了任务类型(朗读与自发说话)和语言熟练度的影响,发现结构化任务精度更高,且在不同熟练度水平间具有良好泛化能力。这些结果提升了自监督模型的可解释性,并展示了其在语音技术应用中的潜力。

原文摘要 · Abstract (English)

SSL speech models capture rich phonetic, prosodic, and acoustic patterns from raw audio, yet how they encode articulatory information across diverse languages remains unclear. Using EMA data from bilingual Finnish-Russian speakers, we evaluate cross-lingual correlations between SSL latent representations and articulatory movements. Models achieve strong prediction performance (Pearson r up to 0.68) even with approximately 5 minutes of training data, with multilingual models outperforming monolingual ones. Intermediate layers encode articulatory features most effectively, and tongue movements are more predictable than lip movements. We also assess the impact of task type (read versus spontaneous speech) and language proficiency, finding higher accuracy for structured tasks and strong generalization across proficiency levels. These results enhance the interpretability of SSL models and show their potential for speech-technology applications.

自监督学习发音建模跨语言EMA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。