arXiv:2506.00375cs.SDeess.AS2025-06被引 2

通过增强伪造痕迹感知提升音频深度伪造检测的泛化能力。

RPRA-ADD: Forgery Trace Enhancement-Driven Audio Deepfake Detection

  • 设计多阶段特征分散增强损失,强化真实与伪造音频的特征差异。
  • 引入聚焦伪造痕迹的注意力机制,动态调整关注区域,提升检测敏感度。
  • 在4个基准数据集上性能超现有方法20%以上,跨域泛化能力强。

现有深度伪造音频检测方法虽有一定效果,但在应对新型伪造技术与演变攻击模式时仍存在泛化能力不足的问题。根源在于模型过度依赖训练数据分布,难以学习到伪造的本质特征边界,且仅使用分类损失难以捕捉真实与伪造音频间的内在差异。为此,本文提出基于重建-感知-强化-注意力的集成框架RPRA-ADD。首先设计全局-局部伪造感知(GLFP)模块,提升对伪造痕迹的声学感知能力;其次提出多阶段分散增强损失(MDEL),在多层级特征空间中实施分散策略,显著强化真实与伪造音频的特征分布差异;此外,引入伪造痕迹聚焦注意力(FTFA)机制,根据重构差异矩阵动态调整注意力权重,提升对伪造片段的关注度。可视化实验表明,该机制不仅增强了对语音段的关注,还提升了模型泛化能力。在ASVspoof2019、ASVspoof2021、CodecFake和FakeSound共4个基准数据集上的实验结果表明,所提方法达到当前最优性能,性能提升超过20%。在涵盖语音、声音与歌唱的3×3严格跨域评估中亦表现优异,展现出强大的跨音频领域泛化能力。

原文摘要 · Abstract (English)

Existing methods for deepfake audio detection have demonstrated some effectiveness. However, they still face challenges in generalizing to new forgery techniques and evolving attack patterns. This limitation mainly arises because the models rely heavily on the distribution of the training data and fail to learn a decision boundary that captures the essential characteristics of forgeries. Additionally, relying solely on a classification loss makes it difficult to capture the intrinsic differences between real and fake audio. In this paper, we propose the RPRA-ADD, an integrated Reconstruction-Perception-Reinforcement-Attention networks based forgery trace enhancement-driven robust audio deepfake detection framework. First, we propose a Global-Local Forgery Perception (GLFP) module for enhancing the acoustic perception capacity of forgery traces. To significantly reinforce the feature space distribution differences between real and fake audio, the Multi-stage Dispersed Enhancement Loss (MDEL) is designed, which implements a dispersal strategy in multi-stage feature spaces. Furthermore, in order to enhance feature awareness towards forgery traces, the Fake Trace Focused Attention (FTFA) mechanism is introduced to adjust attention weights dynamically according to the reconstruction discrepancy matrix. Visualization experiments not only demonstrate that FTFA improves attention to voice segments, but also enhance the generalization capability. Experimental results demonstrate that the proposed method achieves state-of-the-art performance on 4 benchmark datasets, including ASVspoof2019, ASVspoof2021, CodecFake, and FakeSound, achieving over 20% performance improvement. In addition, it outperforms existing methods in rigorous 3*3 cross-domain evaluations across Speech, Sound, and Singing, demonstrating strong generalization capability across diverse audio domains.

音频伪造深度伪造检测注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。