通过多频段注意力机制提升音频伪造检测精度
Audios Don't Lie: Multi-Frequency Channel Attention Mechanism for Audio Deepfake Detection
- 融合多频段注意力与DCT变换,捕捉音频细微频域特征
- 在复杂场景下准确率超传统方法,F1值显著提升
- 适合金融、安防等对语音真实性要求高的场景
随着人工智能技术的快速发展,音频深度伪造技术的应用日益增多,带来了广泛的安 全风险,尤其在金融和社交安全领域,伪造音频的滥用引发严重担忧。为应对这一挑战,本文提出一种基于多频段通道注意力机制(MFCA)与二维离散余弦变换(DCT)的音频深度伪造检测方法。通过将音频信号转换为梅尔频谱图,利用MobileNet V2提取深层特征,并结合MFCA模块对音频信号中不同频段进行加权,有效捕捉音频中的细粒度频域特征,增强对伪造音频的分类能力。实验结果表明,相较于传统方法,该模型在准确率、精确率、召回率及F1分数等指标上均表现显著优势,尤其在复杂音频场景下展现出更强的鲁棒性与泛化能力,为音频深度伪造检测提供了新思路,具有重要实际应用价值。未来将进一步探索更先进的音频检测技术与优化策略,以持续提升检测精度与泛化能力。
原文摘要 · Abstract (English)
With the rapid development of artificial intelligence technology, the application of deepfake technology in the audio field has gradually increased, resulting in a wide range of security risks. Especially in the financial and social security fields, the misuse of deepfake audios has raised serious concerns. To address this challenge, this study proposes an audio deepfake detection method based on multi-frequency channel attention mechanism (MFCA) and 2D discrete cosine transform (DCT). By processing the audio signal into a melspectrogram, using MobileNet V2 to extract deep features, and combining it with the MFCA module to weight different frequency channels in the audio signal, this method can effectively capture the fine-grained frequency domain features in the audio signal and enhance the Classification capability of fake audios. Experimental results show that compared with traditional methods, the model proposed in this study shows significant advantages in accuracy, precision,recall, F1 score and other indicators. Especially in complex audio scenarios, this method shows stronger robustness and generalization capabilities and provides a new idea for audio deepfake detection and has important practical application value. In the future, more advanced audio detection technologies and optimization strategies will be explored to further improve the accuracy and generalization capabilities of audio deepfake detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。