通过跨尺度注意力与一致性学习,实现轻量级高鲁棒性语音伪造检测。
Lightweight Resolution-Aware Audio Deepfake Detection via Cross-Scale Attention and Consistency Learning
- 显式建模多分辨率频谱特征,通过跨尺度注意力对齐不同时间-频率粒度。
- 在真实场景下表现优异:ASVspoof LA 的误报率仅 0.16%,野数据集 AUC 达 0.98。
- 模型仅需 159k 参数、不到 1 GFLOP,适合实际部署,适用于安防与语音验证场景。
语音伪造检测因语音合成与音色转换技术的快速发展而日益困难,尤其在信道失真、重放攻击及真实录音条件下。本文提出一种分辨率感知的音频伪造检测框架,通过跨尺度注意力与一致性学习,显式建模并对齐多分辨率频谱表示。不同于传统单分辨率或隐式特征融合方法,该方法强制不同时间-频率尺度间达成一致。在 ASVspoof 2019(LA 和 PA)、Fake-or-Real(FoR)数据集以及真实环境下音频伪造数据集上进行评估,结果表明:在 ASVspoof LA 上误报率低至 0.16%,在 PA 上为 5.09%,在 FoR 重录音频上为 4.54%,在野数据集中达到 AUC 0.98、EER 4.81%,显著优于单分辨率与非注意力基线。模型轻量高效,仅需 159k 参数,每次推理耗时低于 1 GFLOP,适合实际部署。消融实验证明跨尺度注意力与一致性学习的关键作用,梯度可解释性分析显示模型学习到跨多种伪造条件的一致性语义频谱线索。结果表明,显式跨分辨率建模为下一代音频伪造检测系统提供了原则性、鲁棒且可扩展的基础。
原文摘要 · Abstract (English)
Audio deepfake detection has become increasingly challenging due to rapid advances in speech synthesis and voice conversion technologies, particularly under channel distortions, replay attacks, and real-world recording conditions. This paper proposes a resolution-aware audio deepfake detection framework that explicitly models and aligns multi-resolution spectral representations through cross-scale attention and consistency learning. Unlike conventional single-resolution or implicit feature-fusion approaches, the proposed method enforces agreement across complementary time--frequency scales. The proposed framework is evaluated on three representative benchmarks: ASVspoof 2019 (LA and PA), the Fake-or-Real (FoR) dataset, and the In-the-Wild Audio Deepfake dataset under a speaker-disjoint protocol. The method achieves near-perfect performance on ASVspoof LA (EER 0.16%), strong robustness on ASVspoof PA (EER 5.09%), FoR rerecorded audio (EER 4.54%), and in-the-wild deepfakes (AUC 0.98, EER 4.81%), significantly outperforming single-resolution and non-attention baselines under challenging conditions. The proposed model remains lightweight and efficient, requiring only 159k parameters and less than 1~GFLOP per inference, making it suitable for practical deployment. Comprehensive ablation studies confirm the critical contributions of cross-scale attention and consistency learning, while gradient-based interpretability analysis reveals that the model learns resolution-consistent and semantically meaningful spectral cues across diverse spoofing conditions. These results demonstrate that explicit cross-resolution modeling provides a principled, robust, and scalable foundation for next-generation audio deepfake detection systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。