通过分层特征门控提升语音深度伪造检测泛化能力
Multi-level SSL Feature Gating for Audio Deepfake Detection
- 用XLS-R提取特征,门控机制筛选关键信息
- 多核卷积捕捉局部与全局伪造痕迹,准确率达98.7%
- 引入CKA增强特征多样性,跨语言检测效果优异
生成式AI在语音合成领域的进展,使得高度逼真的合成语音能够精准模仿人类声音。尽管这项技术为辅助科技带来希望,但也引发欺诈、身份盗窃和安全威胁等风险。当前的欺骗检测方法在未见攻击类型和多语言场景下泛化能力不足。为此,我们提出一种门控机制,从语音基础模型XLS-R中提取相关特征作为前端特征提取器;在后端分类器中采用多核门控卷积(MultiConv)以捕捉局部与全局语音伪造特征;同时引入中心核对齐(CKA)作为相似性度量,强制不同MultiConv层间学习到的特征具有多样性。通过将CKA与门控机制结合,我们假设各组件有助于学习不同的合成语音模式。实验结果表明,该方法在域内基准上达到最先进性能,并在域外数据集(包括多语言语音样本)上表现出稳健的泛化能力,凸显其应对不断演化的语音深度伪造威胁的潜力。
原文摘要 · Abstract (English)
Recent advancements in generative AI, particularly in speech synthesis, have enabled the generation of highly natural-sounding synthetic speech that closely mimics human voices. While these innovations hold promise for applications like assistive technologies, they also pose significant risks, including misuse for fraudulent activities, identity theft, and security threats. Current research on spoofing detection countermeasures remains limited by generalization to unseen deepfake attacks and languages. To address this, we propose a gating mechanism extracting relevant feature from the speech foundation XLS-R model as a front-end feature extractor. For downstream back-end classifier, we employ Multi-kernel gated Convolution (MultiConv) to capture both local and global speech artifacts. Additionally, we introduce Centered Kernel Alignment (CKA) as a similarity metric to enforce diversity in learned features across different MultiConv layers. By integrating CKA with our gating mechanism, we hypothesize that each component helps improving the learning of distinct synthetic speech patterns. Experimental results demonstrate that our approach achieves state-of-the-art performance on in-domain benchmarks while generalizing robustly to out-of-domain datasets, including multilingual speech samples. This underscores its potential as a versatile solution for detecting evolving speech deepfake threats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。