arXiv:2606.08038cs.SD2026-06中稿 · Interspeech 2026

数据规模大不等于效果好,多样攻击样本更关键。

Exploring the Scale and Diversity of Speech Anti-spoofing Datasets: Experiments and Analysis

论文配图:Exploring the Scale and Diversity of Speech Anti-spoofing Datasets: Experiments and Analysis
图 1 · 摘自论文原文
  • 拆解数据规模与攻击多样性的影响,对比实验验证。
  • 小而多样的数据集在跨数据集测试中表现显著优于大规模单一数据集。
  • 适合关注语音防欺骗模型泛化能力的研究者阅读。

过去十年间,语音防欺骗数据集规模呈指数级增长,普遍认为更大数据集能带来更好性能。然而,盲目扩大规模是否真正提升模型泛化能力尚不明确。本研究挑战‘以规模为先’的范式,分离训练数据规模与多样性的影响。在代表性数据集上的实验表明:(1)规模并非越大越好;在固定生成方法下过度扩展数据规模,收益微乎其微,甚至因过拟合导致跨域泛化性能下降。(2)多样性胜过规模;包含多种攻击方式的小型复合训练集,在跨数据集评估中显著优于大规模但多样性不足的数据集。结论指出,未来数据集构建应优先考虑生成方法的多样性,以有效提升模型泛化能力。

原文摘要 · Abstract (English)

The scale of speech anti-spoofing datasets has grown exponentially over the past decade, driven by the assumption that larger data leads to better performance. However, it remains unclear whether indiscriminate scaling commensurately improves model generalization. This study challenges the "scale-first" paradigm by decoupling the impacts of training data scale versus diversity. Through experiments on representative datasets, we report two key findings: (1) Larger is not always better. Expanding data scale excessively under fixed generation methods yields negligible returns and may even degrade cross-domain generalization due to overfitting.(2) Diversity outweighs scale. A smaller composite training set featuring diverse attacks significantly outperforms larger-scale datasets with limited diversity in cross-dataset evaluations. We conclude that future dataset construction should prioritize the diversity of generation methods over scale to effectively enhance model generalization.

语音安全数据多样性防欺骗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。