arXiv:2507.04738cs.CLcs.AI2025-07中稿 · Interspeech 2025被引 3

研究自监督语音模型如何学习不同语言的重音特征

Word stress in self-supervised speech models: A cross-linguistic comparison

  • 用Wav2vec 2.0提取语音嵌入,通过分类器判断重音音节
  • 在五种语言中准确区分重读与非重读音节
  • 重音表征具有语言特异性,变重音语言间差异更大

本文研究自监督语音模型(S3M)——特别是Wav2vec 2.0——对五种语言重音表征的学习能力:三种具有可变或词汇重音的语言(荷兰语、英语、德语)和两种具有固定或分隔式重音的语言(匈牙利语、波兰语)。我们在S3M嵌入上训练诊断性重音分类器,结果显示其能在朗读短句中以高准确率区分重读与非重读音节。此外,实验还检验了语言特异性影响,结果表明重音表征具有语言依赖性,可变重音语言组与固定重音语言组之间的差异更为显著。

原文摘要 · Abstract (English)

In this paper we study word stress representations learned by self-supervised speech models (S3M), specifically the Wav2vec 2.0 model. We investigate the S3M representations of word stress for five different languages: Three languages with variable or lexical stress (Dutch, English and German) and two languages with fixed or demarcative stress (Hungarian and Polish). We train diagnostic stress classifiers on S3M embeddings and show that they can distinguish between stressed and unstressed syllables in read-aloud short sentences with high accuracy. We also tested language-specificity effects of S3M word stress. The results indicate that the word stress representations are language-specific, with a greater difference between the set of variable versus the set of fixed stressed languages.

语音模型重音识别跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。