arXiv:2508.07944cs.SDcs.AI2025-08被引 8

构建首个用于分析语音伪造检测偏见的数据集,揭示性别语言年龄等影响检测效果

SCDF: A Speaker Characteristics DeepFake Speech Dataset for Bias Analysis

  • 构建包含23.7万条语音的多语言平衡数据集,覆盖男女、多年龄、五语言
  • 实测发现检测效果在不同性别、语言、年龄群体间差异显著,最高达20%以上
  • 为公平性检测系统开发提供基准,适合研究者和伦理审查人员参考

尽管深度伪造语音检测受到越来越多关注,但其在语音领域中的偏见与公平性问题仍被忽视。为此,我们提出语音特征深度伪造(SCDF)数据集:一个丰富标注的新资源,可用于系统评估深度伪造语音检测中的种族人口学偏见。SCDF包含超过23.7万条语音,涵盖男性与女性说话人、五种语言及广泛年龄范围,实现均衡分布。我们评估了多个先进检测器,发现说话人特征显著影响检测性能,在性别、语言、年龄和合成器类型之间存在明显差异。这些发现凸显了发展具备偏见意识的检测系统的重要性,并为构建符合伦理与监管标准的非歧视性检测系统奠定基础。

原文摘要 · Abstract (English)

Despite growing attention to deepfake speech detection, the aspects of bias and fairness remain underexplored in the speech domain. To address this gap, we introduce the Speaker Characteristics Deepfake (SCDF) dataset: a novel, richly annotated resource enabling systematic evaluation of demographic biases in deepfake speech detection. SCDF contains over 237,000 utterances in a balanced representation of both male and female speakers spanning five languages and a wide age range. We evaluate several state-of-the-art detectors and show that speaker characteristics significantly influence detection performance, revealing disparities across sex, language, age, and synthesizer type. These findings highlight the need for bias-aware development and provide a foundation for building non-discriminatory deepfake detection systems aligned with ethical and regulatory standards.

语音伪造数据集偏见分析公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。