arXiv:2604.08450cs.SDeess.AS2026-04被引 2

DeepFense统一框架提升语音伪造检测可复现性与公平性

DeepFense: A Unified, Modular, and Extensible Framework for Robust Deepfake Audio Detection

  • 构建模块化开源工具链,整合400+模型与100+训练配方
  • 发现预训练特征提取器选择决定性能差异,而非数据质量
  • 揭示高性能模型在音质、性别、语言上的严重偏差

语音深度伪造检测已形成多模型、多数据集、多训练策略的研究体系,但缺乏标准化实现与评估协议,制约了结果的可复现性与跨研究比较。本文提出DeepFense,一个全面的开源PyTorch工具包,集成最新架构、损失函数与增强管道,包含100多个训练配方。利用该框架,我们对400多个模型进行了大规模评估。结果表明,精心筛选的训练数据虽有助于跨领域泛化,但预训练前端特征提取器的选择才是整体性能波动的主导因素。关键发现是,高性能模型在音频质量、说话人性别和语言上存在严重偏差。DeepFense将为实际部署提供必要工具,支持公平的数据选择与前端微调。

原文摘要 · Abstract (English)

Speech deepfake detection is a well-established research field with different models, datasets, and training strategies. However, the lack of standardized implementations and evaluation protocols limits reproducibility, benchmarking, and comparison across studies. In this work, we present DeepFense, a comprehensive, open-source PyTorch toolkit integrating the latest architectures, loss functions, and augmentation pipelines, alongside over 100 recipes. Using DeepFense, we conducted a large-scale evaluation of more than 400 models. Our findings reveal that while carefully curated training data improves cross-domain generalization, the choice of pre-trained front-end feature extractor dominates overall performance variance. Crucially, we show severe biases in high-performing models regarding audio quality, speaker gender, and language. DeepFense is expected to facilitate real-world deployment with the necessary tools to address equitable training data selection and front-end fine-tuning.

语音伪造检测深度学习公平性工具框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。