通过分析音频帧级潜在信息熵,提升语音伪造检测的泛化能力。
Generalized Audio Deepfake Detection Using Frame-level Latent Information Entropy
- 从信息论角度提取音频帧级潜在表示的信息熵差异
- 在最新合成数据集上达到顶尖检测性能,泛化能力显著
- 适合关注语音伪造检测泛化与可解释性的研究者
通用性是音频深度伪造检测的关键,因文本转语音(TTS)和语音转换(VC)技术快速演进。本文提出一种新思路:利用内在差异区分真实与伪造音频。从信息论视角看,真实音频蕴含密集高信息量特征,而伪造音频源于低维、信息较少的表示。为此,我们提出帧级潜在信息熵检测器(f-InfoED),从帧级潜在表示中提取独特信息熵以识别深度伪造。同时引入AdaLAM,通过可训练适配器增强预训练音频模型的特征提取能力。为全面评估,构建了基于最新TTS与VC方法的音频深度伪造取证2024(ADFF 2024)数据集。大量实验表明,所提方法达当前最优性能,且具备卓越泛化能力。分析进一步验证了AdaLAM在提取判别性特征及f-InfoED在利用潜在熵信息实现更通用检测方面的有效性。
原文摘要 · Abstract (English)
Generalizability, the capacity of a robust model to perform effectively on unseen data, is crucial for audio deepfake detection due to the rapid evolution of text-to-speech (TTS) and voice conversion (VC) technologies. A promising approach to differentiate between bonafide and spoof samples lies in identifying intrinsic disparities to enhance model generalizability. From an information-theoretic perspective, we hypothesize the information content is one of the intrinsic differences: bonafide sample represents a dense, information-rich sampling of the real world, whereas spoof sample is typically derived from lower-dimensional, less informative representations. To implement this, we introduce frame-level latent information entropy detector(f-InfoED), a framework that extracts distinctive information entropy from latent representations at the frame level to identify audio deepfakes. Furthermore, we present AdaLAM, which extends large pre-trained audio models with trainable adapters for enhanced feature extraction. To facilitate comprehensive evaluation, the audio deepfake forensics 2024 (ADFF 2024) dataset was built by the latest TTS and VC methods. Extensive experiments demonstrate that our proposed approach achieves state-of-the-art performance and exhibits remarkable generalization capabilities. Further analytical studies confirms the efficacy of AdaLAM in extracting discriminative audio features and f-InfoED in leveraging latent entropy information for more generalized deepfake detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。