提出首个应对通信降质的语音伪造检测框架,提升真实场景下识别能力。
Multi-Granularity Adaptive Time-Frequency Attention Framework for Audio Deepfake Detection under Real-World Communication Degradations
- 设计多粒度自适应注意力机制,动态捕捉时频特征
- 在六种编码器五级丢包下均超越现有方法
- 适合实际通信环境中的语音安全防护应用
合成语音的兴起对音频通信构成日益严重的威胁。尽管现有语音伪造检测(ADD)方法在干净条件下表现良好,但在真实通信环境中常见的分组丢失和语音编码压缩等降质情况下,性能显著下降。本文提出首个统一的鲁棒性ADD框架,可有效适配多种时频(TF)表示。核心是新颖的多粒度自适应注意力(MGAA)架构,采用可配置的多尺度注意力头,捕获不同粒度下的全局与局部感受野。随后的自适应融合机制根据时频区域显著性调整并融合各注意力分支,使模型能动态重分配关注点,有效定位并放大细微伪造痕迹。大量实验表明,该框架在六种语音编码器及五级分组丢失的各种真实降质场景下,持续优于最先进基线。对比分析显示,MGAA增强特征显著提升了真实与伪造音频类别的可分性,并锐化决策边界。结果凸显了该框架在真实通信环境中的鲁棒性与实用部署潜力。
原文摘要 · Abstract (English)
The rise of highly convincing synthetic speech poses a growing threat to audio communications. Although existing Audio Deepfake Detection (ADD) methods have demonstrated good performance under clean conditions, their effectiveness drops significantly under degradations such as packet losses and speech codec compression in real-world communication environments. In this work, we propose the first unified framework for robust ADD under such degradations, which is designed to effectively accommodate multiple types of Time-Frequency (TF) representations. The core of our framework is a novel Multi-Granularity Adaptive Attention (MGAA) architecture, which employs a set of customizable multi-scale attention heads to capture both global and local receptive fields across varying TF granularities. A novel adaptive fusion mechanism subsequently adjusts and fuses these attention branches based on the saliency of TF regions, allowing the model to dynamically reallocate its focus according to the characteristics of the degradation. This enables the effective localization and amplification of subtle forgery traces. Extensive experiments demonstrate that the proposed framework consistently outperforms state-of-the-art baselines across various real-world communication degradation scenarios, including six speech codecs and five levels of packet losses. In addition, comparative analysis reveals that the MGAA-enhanced features significantly improve separability between real and fake audio classes and sharpen decision boundaries. These results highlight the robustness and practical deployment potential of our framework in real-world communication environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。