arXiv:2509.23618cs.SDcs.AI2025-09被引 3

提出新方法提升语音深度伪造检测的泛化能力。

Generalizable Speech Deepfake Detection via Information Bottleneck Enhanced Adversarial Alignment

  • 通过自适应对抗对齐抑制攻击特异性伪影
  • 信息瓶颈机制去除干扰因素,保留可迁移特征
  • 在多个数据集上达到当前最佳效果,适合实际部署

神经语音合成技术已能生成高度逼真的语音深度伪造,带来严重安全风险。由于伪造方法、说话人、信道和录制条件之间的分布差异,语音深度伪造检测面临挑战。本文探索学习共享判别特征以提升鲁棒性,提出信息瓶颈增强的置信度感知对抗网络(IB-CAAN)。置信度引导的对抗对齐可自适应抑制攻击特异性伪影,同时不破坏判别线索;信息瓶颈则移除冗余变化,保留可迁移特征。在ASVspoof 2019/2021、ASVspoof 5和In-the-Wild数据集上的实验表明,IB-CAAN始终优于基线模型,并在多个基准上达到最先进性能。

原文摘要 · Abstract (English)

Neural speech synthesis techniques have enabled highly realistic speech deepfakes, posing major security risks. Speech deepfake detection is challenging due to distribution shifts across spoofing methods and variability in speakers, channels, and recording conditions. We explore learning shared discriminative features as a path to robust detection and propose Information Bottleneck enhanced Confidence-Aware Adversarial Network (IB-CAAN). Confidence-guided adversarial alignment adaptively suppresses attack-specific artifacts without erasing discriminative cues, while the information bottleneck removes nuisance variability to preserve transferable features. Experiments on ASVspoof 2019/2021, ASVspoof 5, and In-the-Wild demonstrate that IB-CAAN consistently outperforms baseline and achieves state-of-the-art performance on many benchmarks.

语音伪造对抗学习特征提取泛化检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。