arXiv:2506.08981cs.CL2025-06中稿 · Interspeech 2025被引 3

构建芬兰语与俄语双语发音数据集,研究母语、第二语言及模仿口音的语音差异。

FROST-EMA: Finnish and Russian Oral Speech Dataset of Electromagnetic Articulography Measurements with L1, L2 and Imitated L2 Accents

  • 采集18名双语者在母语、第二语言及模仿口音下的电磁发音测量数据。
  • 发现第二语言和模仿口音显著影响自动说话人验证系统的识别准确率。
  • 适用于语音学、言语技术及跨语言语音研究者。

本文介绍一个新的FROST-EMA(芬兰语与俄语口腔发音电磁测量数据集)语料库,包含18名双语说话者在母语(L1)、第二语言(L2)及模仿的第二语言(假外语音)中的发音数据。该数据集支持从语音学与技术角度研究语言变异性。为此,我们进行了两项初步案例研究:第一项分析了第二语言和模仿口音对自动说话人验证系统性能的影响;第二项展示了某位说话者在母语、第二语言及假口音状态下的发音器官运动模式。

原文摘要 · Abstract (English)

We introduce a new FROST-EMA (Finnish and Russian Oral Speech Dataset of Electromagnetic Articulography) corpus. It consists of 18 bilingual speakers, who produced speech in their native language (L1), second language (L2), and imitated L2 (fake foreign accent). The new corpus enables research into language variability from phonetic and technological points of view. Accordingly, we include two preliminary case studies to demonstrate both perspectives. The first case study explores the impact of L2 and imitated L2 on the performance of an automatic speaker verification system, while the second illustrates the articulatory patterns of one speaker in L1, L2, and a fake accent.

语音学发音数据集跨语言语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。