arXiv:2603.09007cs.SDcs.AI2026-03中稿 · IEEE CAI Conferenc…被引 5

分析语音伪造检测中性别偏差,发现主流指标掩盖了真实不公平性。

Gender Fairness in Audio Deepfake Detection: Performance and Disparity Analysis

  • 用四个音频特征训练ResNet-18,对比AASIST模型性能。
  • 整体误报率差异小,但性别间错误分布不均,存在隐蔽偏差。
  • 强调需引入公平性度量,适合关注模型公正性的研究者。

语音伪造检测旨在识别由人工智能生成的语音与真实人声,已成为语音生物识别系统中的关键问题。随着合成语音质量不断提升,其被用于身份盗用和冒名顶替等非法行为的风险增加。尽管近年来该领域取得显著进展,但性别偏见问题仍处于探索初期。本文对音频伪造检测模型中的性别相关性能与公平性进行了全面分析。基于ASVspoof 5数据集,训练了ResNet-18分类器,并在四种不同音频特征上评估检测性能,同时与基线AASIST模型进行比较。除传统指标如等错误率(EER%)外,还引入五种成熟公平性度量来量化性别差异。结果表明,即使总体EER性别差异较小,公平性视角下仍揭示出被聚合指标掩盖的错误分布不均现象。这说明仅依赖标准指标不可靠,而公平性度量能提供关于特定人群失效模式的关键洞察。本研究强调了在构建更公平、鲁棒且可信的语音伪造检测系统时,开展公平性评估的重要性。

原文摘要 · Abstract (English)

Audio deepfake detection aims to detect real human voices from those generated by Artificial Intelligence (AI) and has emerged as a significant problem in the field of voice biometrics systems. With the ever-improving quality of synthetic voice, the probability of such a voice being exploited for illicit practices like identity thest and impersonation increases. Although significant progress has been made in the field of Audio Deepfake Detection in recent times, the issue of gender bias remains underexplored and in its nascent stage In this paper, we have attempted a thorough analysis of gender dependent performance and fairness in audio deepfake detection models. We have used the ASVspoof 5 dataset and train a ResNet-18 classifier and evaluate detection performance across four different audio features, and compared the performance with baseline AASIST model. Beyond conventional metrics such as Equal Error Rate (EER %), we incorporated five established fairness metrics to quantify gender disparities in the model. Our results show that even when the overall EER difference between genders appears low, fairness-aware evaluation reveals disparities in error distribution that are obscured by aggregate performance measures. These findings demonstrate that reliance on standard metrics is unreliable, whereas fairness metrics provide critical insights into demographic-specific failure modes. This work highlights the importance of fairness-aware evaluation for developing a more equitable, robust, and trustworthy audio deepfake detection system.

语音伪造公平性深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。