构建大规模合成语音数据集,测试模型在分布偏移下的泛化能力。
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts
- 设计可控的分布偏移实验,覆盖多语言、多语音合成系统。
- 7个源域、3000+小时语音,6种TTS与12种声码器生成数据。
- 现有检测模型在分布偏移下性能显著下降,需谨慎部署。
合成语音检测近年来取得显著进展,现有方法在多个基准上表现良好。然而,学术基准上的低误差率能否真实反映实际场景?现实中,训练数据固定,测试时却可能因说话人特征、情感表达、语言、声学条件变化及新合成方法出现而产生分布偏移。尽管已有数据集涵盖部分偏移类型,但跨数据集源数据与合成系统不一致,限制了系统性分析。加之文本转语音(TTS)和声码器技术快速演进,合成语音多样性持续增加。为此,我们提出ShiftySpeech,一个包含超过3000小时合成语音的大规模基准,覆盖7个源域、6种TTS系统、12种声码器及3种语言。该数据集专为评估模型在受控分布偏移下的泛化能力而设计,全面覆盖现代合成语音生成技术。结果表明,基于自监督特征的先进检测方法在所有分布偏移下性能均显著下降。研究提示:生产环境中使用合成语音检测方法前,必须评估其对预期分布偏移的鲁棒性。
原文摘要 · Abstract (English)
The problem of synthetic speech detection has enjoyed considerable attention, with recent methods achieving low error rates across several established benchmarks. However, to what extent can low error rates on academic benchmarks translate to more realistic conditions? In practice, while the training set is fixed at one point in time, test-time conditions may exhibit distribution shifts relative to the training conditions, such as changes in speaker characteristics, emotional expressiveness, language and acoustic conditions, and the emergence of novel synthesis methods. Although some existing datasets target subsets of these distribution shifts, systematic analysis remains difficult due to inconsistencies between source data and synthesis systems across datasets. This difficulty is further exacerbated by the rapid development of new text-to-speech (TTS) and vocoder systems, which continually expand the diversity of synthetic speech. To enable systematic benchmarking of model performance under distribution shifts, we introduce ShiftySpeech, a large-scale benchmark comprising over 3,000 hours of synthetic speech across 7 source domains, 6 TTS systems, 12 vocoders, and 3 languages. ShiftySpeech is specifically designed to evaluate model generalization under controlled distribution shifts while ensuring broad coverage of modern synthetic speech generation techniques. It fills a key gap in current benchmarks by supporting fine-grained, controlled analysis of generalization robustness. All tested distribution shifts significantly degrade detection performance of state-of-the-art detection approaches based on self-supervised features. Overall, our findings suggest that reliance on synthetic speech detection methods in production environments should be carefully evaluated based on anticipated distribution shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。