arXiv:2603.02364cs.SDeess.AS2026-03中稿 · Interspeech 2026被引 2

跨66种语言测试语音伪造检测,发现语言差异是重要干扰因素

When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus

  • 构建66语言伪造语音数据集,覆盖45种低资源语言
  • 在跨语言场景下,检测性能波动显著,语言成独立域偏移源
  • 无需目标语言真声样本,通过阈值迁移评估模型鲁棒性

我们提出LRLspoof,一个大规模多语言合成语音语料库,用于跨语言伪造检测。该语料库包含2,732小时音频,由24个开源文本转语音系统生成,覆盖66种语言,其中包括45种符合我们定义的低资源语言。为评估模型鲁棒性而不依赖目标领域真实语音,我们采用阈值迁移策略:对每个模型在合并的外部基准上校准等错误率(EER)操作点,然后应用所得阈值,报告欺骗拒绝率(SRR)。结果表明,模型在跨语言检测中存在显著差异,即使在受控条件下,不同语言间的欺骗拒绝率也变化明显,凸显语言本身是伪造检测中的独立域偏移来源。数据集已公开于HuggingFace与ModelScope。

原文摘要 · Abstract (English)

We introduce LRLspoof, a large-scale multilingual synthetic-speech corpus for cross-lingual spoof detection, comprising 2,732 hours of audio generated with 24 open-source TTS systems across 66 languages, including 45 low-resource languages under our operational definition. To evaluate robustness without requiring target-domain bonafide speech, we benchmark 11 publicly available countermeasures using threshold transfer: for each model we calibrate an EER operating point on pooled external benchmarks and apply the resulting threshold, reporting spoof rejection rate (SRR). Results show model-dependent cross-lingual disparity, with spoof rejection varying markedly across languages even under controlled conditions, highlighting language as an independent source of domain shift in spoof detection. The dataset is publicly available at \href{https://huggingface.co/datasets/lab260/LRLspoof}{\textbf{\underline{\textit{HuggingFace}}}} and \href{https://modelscope.cn/datasets/lab260/LRLspoof}{\textbf{\underline{\textit{ModelScope}}}}

语音伪造跨语言低资源语言检测评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。