arXiv:2601.21960eess.AScs.SD2026-01被引 4

挑战跨语言语音识别性能下降,推动更公平的多语种声纹验证技术。

TidyVoice 2026 Challenge Evaluation Plan

  • 基于40多种语言的TidyVoiceX数据集,专设语言切换场景测试系统鲁棒性。
  • 采用跨语言测试的等错误率(EER)作为核心评估指标,突出语言不匹配问题。
  • 适合关注多语种语音识别、公平性与可扩展声纹验证的研究者。

语音验证系统在语言不匹配场景下性能显著下降,这一挑战因领域长期依赖英语数据而加剧。为此,我们提出针对跨语言声纹验证的TidyVoice挑战。该挑战基于新型TidyVoice基准中的TidyVoiceX数据集,这是一个大规模多语言语料库,源自Mozilla Common Voice,特别设计以隔离约40种语言间语言切换的影响。参赛者需构建对语言不匹配具有鲁棒性的系统,主要通过跨语言测试的等错误率(EER)进行评估。通过提供标准化数据、开源基线和严格评估协议,本挑战旨在推动研究向更公平、更包容、语言无关的声纹识别技术发展,直接呼应Interspeech 2026主题“Speaking Together”。

原文摘要 · Abstract (English)

The performance of speaker verification systems degrades significantly under language mismatch, a critical challenge exacerbated by the field's reliance on English-centric data. To address this, we propose the TidyVoice Challenge for cross-lingual speaker verification. The challenge leverages the TidyVoiceX dataset from the novel TidyVoice benchmark, a large-scale, multilingual corpus derived from Mozilla Common Voice, and specifically curated to isolate the effect of language switching across approximately 40 languages. Participants will be tasked with building systems robust to this mismatch, with performance primarily evaluated using the Equal Error Rate on cross-language trials. By providing standardized data, open-source baselines, and a rigorous evaluation protocol, this challenge aims to drive research towards fairer, more inclusive, and language-independent speaker recognition technologies, directly aligning with the Interspeech 2026 theme, "Speaking Together."

声纹验证跨语言多语种公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。