arXiv:2607.01161eess.AScs.CL2026-07

通过同人跨语言测试,发现语音识别性能下降主因是语言差异而非说话人变化。

Disentangling Speaker and Language Effects in Cross-Lingual Speaker Verification for Iberian Languages

论文配图:Disentangling Speaker and Language Effects in Cross-Lingual Speaker Verification for Iberian Languages
图 1 · 摘自论文原文
  • 构建五种伊比利亚语同说话人跨语言测试集,固定说话人身份。
  • 语言不匹配导致性能下降,说话人差异仅占部分影响。
  • 揭示跨语言语音验证中语言依赖的真正根源,适合语音系统研究者。

跨语言语音验证(SV)系统在注册与测试语句使用不同语言时通常表现下降。然而,标准评估协议将语言不匹配与说话人差异混淆,因跨语言评估通常涉及不同说话人。本文为五种伊比利亚语言构建了同说话人跨语言评估集,实现说话人身份恒定下的跨语言SV分析。我们在此设置下应用此前表现出强语言依赖性的HuBERT-based SV系统,并利用跨语言迁移矩阵(CLTM)分析成对跨语言迁移。结果表明,说话人相关变异性贡献了部分性能损失,但语言不匹配仍是跨语言性能下降的主要原因。该研究为跨语言语音验证中的语言依赖性提供了更精确的刻画。

原文摘要 · Abstract (English)

Cross-lingual speaker verification (SV) systems typically exhibit performance degradation when enrollment and test utterances are spoken in different languages. However, standard evaluation protocols confound language mismatch with inter-speaker variability, as evaluation is generally performed with different speakers across languages. In this work, we introduce a bilingual same-speaker evaluation set for five Iberian languages, enabling analysis of cross-lingual SV under constant speaker identity. We apply this setup to a HuBERT-based SV system previously shown to exhibit strong language dependence, and analyze results using the Cross-Lingual Transfer Matrix (CLTM) to study pairwise cross-lingual transfer. Our results show that speaker-related variability accounts for part of the observed degradation, but language mismatch remains the main driver of cross-lingual performance loss. These findings provide a more precise characterization of language dependence in cross-lingual SV.

语音验证跨语言说话人分离伊比利亚语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。