arXiv:2506.01156cs.CLcs.SD2025-06中稿 · Interspeech 2025 c…被引 1

用少量外语数据训练芬兰瑞典语发音纠错模型

Mispronunciation Detection Without L2 Pronunciation Dataset in Low-Resource Setting: A Case Study in Finland Swedish

  • 仅用母语者语音训练,通过熵正则化适配发音检测
  • 在33分钟外语语音上实现43.2%召回率与29.8%精确率
  • 方法简洁且可迁移,适合其他低资源语言

发音纠错(MD)模型是众多语言学习应用的核心。然而,多数系统针对英语等主流语言构建,而像芬兰瑞典语(FS)这样的低资源语言缺乏此类工具。本文提出针对FS的MD模型,使用89小时母语者自发语音训练,测试数据为33分钟非母语者朗读转录语音。采用带熵正则化的多语言wav2vec 2.0模型,推理后引入温度缩放和top-k归一化以提升适配性。该方法主要创新在于其简洁性,所需外语数据极少,且具备语言无关性,适用于其他低资源语言。相比基线模型(召回率77.5%,精确率17.6%),本方法实现43.2%召回率与29.8%精确率,有效平衡两者性能。

原文摘要 · Abstract (English)

Mispronunciation detection (MD) models are the cornerstones of many language learning applications. Unfortunately, most systems are built for English and other major languages, while low-resourced language varieties, such as Finland Swedish (FS), lack such tools. In this paper, we introduce our MD model for FS, trained on 89 hours of first language (L1) speakers' spontaneous speech and tested on 33 minutes of L2 transcribed read-aloud speech. We trained a multilingual wav2vec 2.0 model with entropy regularization, followed by temperature scaling and top-k normalization after the inference to better adapt it for MD. The main novelty of our method lies in its simplicity, requiring minimal L2 data. The process is also language-independent, making it suitable for other low-resource languages. Our proposed algorithm allows us to balance Recall (43.2%) and Precision (29.8%), compared with the baseline model's Recall (77.5%) and Precision (17.6%).

发音纠错低资源语言wav2vec 2.0

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。