构建了21万+单语和4500+多语说话人验证数据集,解决跨语言识别难题
TidyVoice: A Curated Multilingual Dataset for Speaker Verification Derived from Common Voice
- 从Common Voice清理出高一致性说话人数据,按语言分单语/多语两类
- 在21.2万单语数据上微调模型,达到0.35%错误率(EER)
- 提升模型泛化能力,在未见对话数据上表现更优,适合语音安全研究
稳健的多语言说话人识别系统受限于缺乏大规模、公开且多语言的数据集,尤其在读话语音场景下更为明显。为此,我们基于Mozilla Common Voice语料库,通过缓解其客户端ID带来的说话人异质性,构建了TidyVoice数据集。当前,TidyVoice包含超过212,000名单语说话人(Tidy-M)和约4,500名多语说话人(Tidy-X)的数据,分为两种测试条件:Tidy-M包含81种语言中单语说话人的目标与非目标试次;Tidy-X包含多语说话人在同语言与跨语言试次中的目标与非目标试次。我们采用两种ResNet架构,在完整的Tidy-M分区上微调,实现0.35%的EER。此外,该微调显著提升了模型泛化能力,在未见过的对话访谈数据(来自CANDOR语料库)上表现更好。完整数据集、评估试次及模型已公开发布,为社区提供新资源。
原文摘要 · Abstract (English)
The development of robust, multilingual speaker recognition systems is hindered by a lack of large-scale, publicly available and multilingual datasets, particularly for the read-speech style crucial for applications like anti-spoofing. To address this gap, we introduce the TidyVoice dataset derived from the Mozilla Common Voice corpus after mitigating its inherent speaker heterogeneity within the provided client IDs. TidyVoice currently contains training and test data from over 212,000 monolingual speakers (Tidy-M) and around 4,500 multilingual speakers (Tidy-X) from which we derive two distinct conditions. The Tidy-M condition contains target and non-target trials from monolingual speakers across 81 languages. The Tidy-X condition contains target and non-target trials from multilingual speakers in both same- and cross-language trials. We employ two architectures of ResNet models, achieving a 0.35% EER by fine-tuning on our comprehensive Tidy-M partition. Moreover, we show that this fine-tuning enhances the model's generalization, improving performance on unseen conversational interview data from the CANDOR corpus. The complete dataset, evaluation trials, and our models are publicly released to provide a new resource for the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。