arXiv:2607.22923cs.CL2026-07

用简单语言归一化提升跨语言语音验证性能。

Simple Language Normalization Wins: Cross-Lingual Speaker Verification for the TidyVoice 2026 Challenge

论文配图:Simple Language Normalization Wins: Cross-Lingual Speaker Verification for the TidyVoice 2026 Challenge
图 1 · 摘自论文原文
  • 在嵌入空间投影去除语言干扰,实现简洁有效的语言归一化。
  • 开发集误报率降至2.18%,评测得分达8.40,优于复杂模型。
  • 适合关注跨语言语音验证的工业与研究场景。

跨语言不匹配仍是现代说话人验证中的主要性能瓶颈。TidyVoice2026挑战赛聚焦无文本依赖的跨语言验证,包含40种语言中3,666名训练者和808名开发集说话人,以及38种未见语言中的2,200名评测说话人,测试时无语言标签。基于在VoxBlink2和VoxCeleb2上预训练、并在TidyVoice上微调的SimAM-ResNet34基线,本文重新审视了干扰属性投影(NAP)作为嵌入空间中的简单语言归一化步骤。通过跨语言同说话人差异估计出紧凑语言子空间,并将嵌入投影到其正交补空间后进行余弦打分,结合自适应对称归一化(AS-Norm)。该方法使开发集等错误率从余弦法的2.97%、AS-Norm的2.70%降至2.18%,在Codabench评测中取得8.40分,表明简单的后端语言归一化可媲美更复杂的系统。

原文摘要 · Abstract (English)

Cross-lingual mismatch remains a key source of overall degradation in modern speaker verification. The TidyVoice2026 Challenge targets this setting with text-independent verification, comprising 3,666 training and 808 development speakers in 40 languages and 2,200 evaluation speakers in 38 unseen languages, without language labels at test time. Starting from the official SimAM-ResNet34 baseline pretrained on VoxBlink2 and VoxCeleb2 and fine-tuned on TidyVoice, we revisit Nuisance Attribute Projection (NAP) as a simple language-normalization step in the embedding space. We estimate a compact language subspace from cross-language same-speaker differences and project embeddings onto its orthogonal complement before cosine scoring with Adaptive Symmetric score normalization. This reduces development EER from 2.97\% with cosine and 2.70\% with AS-Norm to 2.18\% and yields a Codabench evaluation score of 8.40, showing that simple back-end language normalization can rival more complex systems.

语音验证跨语言嵌入归一化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。