arXiv:2409.08346eess.AScs.AI2024-09中稿 · the IEEE Spoken La…被引 15

用语音合成扩充数据,让英语训练的反欺骗模型更好应对其他语言。

Towards Quantifying and Reducing Language Mismatch Effects in Cross-Lingual Speech Anti-Spoofing

  • 通过语音合成引入多语言特征,增强单语模型跨语言能力。
  • 在12种语言、超300万样本上测试,性能下降减少超15%。
  • 适合多语言和低资源语言场景,实现简单且效果显著。

语言差异对语音反欺骗系统的影响显著,但相关研究与量化仍有限。现有反欺骗数据集以英语为主,多语言数据收集成本高,阻碍了语言无关模型的训练。本文评估了在英语数据上训练、却在其他语言上测试的顶尖反欺骗系统,发现性能明显下降。为此提出一种创新方法:基于发音特征的语音合成数据扩充(ACCENT),通过引入多样化的语言知识,提升单语模型的跨语言适应性。实验在包含超过300万样本的大规模数据集上进行,其中训练样本180万,测试样本近120万,覆盖12种语言。初步量化了语言不匹配效应,并通过ACCENT方法将其显著降低超过15%。该方法实现简单,在多语言及低资源语言场景中具有广阔应用前景。

原文摘要 · Abstract (English)

The effects of language mismatch impact speech anti-spoofing systems, while investigations and quantification of these effects remain limited. Existing anti-spoofing datasets are mainly in English, and the high cost of acquiring multilingual datasets hinders training language-independent models. We initiate this work by evaluating top-performing speech anti-spoofing systems that are trained on English data but tested on other languages, observing notable performance declines. We propose an innovative approach - Accent-based data expansion via TTS (ACCENT), which introduces diverse linguistic knowledge to monolingual-trained models, improving their cross-lingual capabilities. We conduct experiments on a large-scale dataset consisting of over 3 million samples, including 1.8 million training samples and nearly 1.2 million testing samples across 12 languages. The language mismatch effects are preliminarily quantified and remarkably reduced over 15% by applying the proposed ACCENT. This easily implementable method shows promise for multilingual and low-resource language scenarios.

语音反欺骗跨语言数据增强TTS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。