39个深度伪造语音数据集审计揭示公平性评估难、数据重复率高问题。
Ethical and Technical Limits of Deepfake Speech Datasets
- 系统梳理39个数据集,分析可访问性、标注完整性与真实语音来源
- 90%以上数据集缺人口统计信息,无法开展公平性子组分析
- 多个数据集共享相同真实语音源,影响跨数据集评测可信度
深度伪造语音检测系统的鲁棒性与公平性评估,其可信度取决于所用数据集的质量。我们对当前深度伪造语音领域进行了数据集层面的审计,共收集并分析了39个数据集,涵盖可访问性、文档完整性、人口与语言覆盖范围、数据规模以及真实语音来源等关键属性。研究发现:第一,公平性评估基本不可行,因多数数据集缺乏人口统计元数据,仅有少数包含性别或语言标签,导致无法进行有意义的子群体分析;第二,多个数据集存在显著的真实语音来源重叠,可能削弱跨数据集评估的有效性,导致模型泛化能力被过度夸大。
原文摘要 · Abstract (English)
Claims about the robustness and fairness of deepfake speech detectors are only as credible as the datasets used to train and evaluate those systems. We present a dataset-level audit of the deepfake speech landscape. We compile and analyze 39 deepfake speech datasets, examining key attributes including accessibility, documentation, demographic and language coverage, dataset scale, and the underlying bona fide speech sources. Our audit reveals two important takeaways. Firstly, fairness assessment is largely infeasible because most datasets lack demographic metadata, and only a few contain gender or language labels. This prevents any meaningful subgroup analysis and leaves other demographic attributes unaddressed. Secondly, we identify substantial overlap in underlying bona fide source corpora across datasets, which can undermine cross-dataset evaluation and lead to overstated generalization claims.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。