通过显式建模相位信息提升深度伪造检测精度
Phase4DFD: Multi-Domain Phase-Aware Attention for Deepfake Detection
- 引入相位感知注意力机制,捕捉合成生成带来的相位不连续性
- 在CIFAKE和DFFD数据集上超越现有方法,准确率显著提升
- 适合关注频率域特征与伪造检测的算法研究者
近期深度伪造检测方法越来越多地利用频域表示来揭示空间域难以察觉的篡改痕迹。然而,现有方法主要依赖频谱幅度,隐式忽略了相位信息的作用。本文提出Phase4DFD,一种相位感知的频域深度伪造检测框架,通过可学习的注意力机制显式建模相位与幅度的交互关系。该方法在标准RGB输入基础上,融合快速傅里叶变换(FFT)幅度和局部二值模式(LBP)表示,以暴露仅靠空间分析无法识别的细微合成痕迹。关键在于,我们设计了输入级相位感知注意力模块,利用合成生成常引入的相位不连续性,引导模型在主干特征提取前聚焦于最具操纵指示性的频域模式。经由高效的BNext M主干网络处理,辅以可选的通道-空间注意力进行语义特征优化。在CIFAKE和DFFD数据集上的大量实验表明,Phase4DFD在保持低计算开销的同时,优于当前最先进的空间与频域检测器。全面的消融实验进一步证实,显式相位建模提供了与幅度信息互补且非冗余的判别性信息。
原文摘要 · Abstract (English)
Recent deepfake detection methods have increasingly explored frequency domain representations to reveal manipulation artifacts that are difficult to detect in the spatial domain. However, most existing approaches rely primarily on spectral magnitude, implicitly under exploring the role of phase information. In this work, we propose Phase4DFD, a phase aware frequency domain deepfake detection framework that explicitly models phase magnitude interactions via a learnable attention mechanism. Our approach augments standard RGB input with Fast Fourier Transform (FFT) magnitude and local binary pattern (LBP) representations to expose subtle synthesis artifacts that remain indistinguishable under spatial analysis alone. Crucially, we introduce an input level phase aware attention module that uses phase discontinuities commonly introduced by synthetic generation to guide the model toward frequency patterns that are most indicative of manipulation before backbone feature extraction. The attended multi domain representation is processed by an efficient BNext M backbone, with optional channel spatial attention applied for semantic feature refinement. Extensive experiments on the CIFAKE and DFFD datasets demonstrate that our proposed model Phase4DFD outperforms state of the art spatial and frequency-based detectors while maintaining low computational overhead. Comprehensive ablation studies further confirm that explicit phase modeling provides complementary and non-redundant information beyond magnitude-only frequency representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。