arXiv:2604.26465cs.SD2026-04

用扩散模型生成难例,提升语音深度伪造检测的泛化能力

Diffusion Reconstruction towards Generalizable Audio Deepfake Detection

论文配图:Diffusion Reconstruction towards Generalizable Audio Deepfake Detection
图 1 · 摘自论文原文
  • 通过扩散模型生成难以分辨的伪造音频样本
  • 在未见攻击下平均错误率降低,优于基线模型
  • 适合需要强泛化能力的语音安全检测场景

在语音深度伪造检测(ADD)中,应对未知攻击的鲁棒泛化仍面临挑战,主要源于生成模型的快速演进。为此,我们提出以难例分类为核心的框架:能够区分困难样本的模型自然也能有效处理简单样本。我们研究了多种重建范式,发现基于扩散的方法最适于生成难例。此外,通过多层特征聚合和引入正则化辅助对比学习(RACL)目标,进一步提升了模型泛化能力。实验表明,所提方法在未见攻击下表现出显著更强的泛化性能,最佳模型相较基线实现平均等错误率(EER)的大幅下降。

原文摘要 · Abstract (English)

Achieving robust generalization against unseen attacks remains a challenge in Audio Deepfake Detection (ADD), driven by the rapid evolution of generative models. To address this, we propose a framework centered on hard sample classification. The core idea is that a model capable of distinguishing challenging hard samples is inherently equipped to handle simpler cases effectively. We investigate multiple reconstruction paradigms, identifying the diffusion-based method as optimal for generating hard samples. Furthermore, we leverage multi-layer feature aggregation and introduce a Regularization-Assisted Contrastive Learning (RACL) objective to enhance generalizability. Experiments demonstrate the superior generalization of our approach, with our best model achieving a significant reduction in the average Equal Error Rate (EER) compared to the baseline.

语音生成深度伪造扩散模型检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。