通过空间与频域融合,提升伪造图像的细微特征识别能力。
D2Fusion: Dual-domain Fusion with Feature Superposition for Deepfake Detection
- 设计双向注意力捕捉局部伪造痕迹,定位更精准
- 引入频域细粒度注意力,提取全局细微伪造信息
- 提出特征叠加策略,增强真实与伪造特征差异
深度伪造检测对遏制其社会危害至关重要。然而,现有方法因缺乏域间内在交互,未能充分挖掘不同域中的伪造特征。本文提出一种双域融合框架,首先在空间域使用双向注意力模块捕捉局部伪造位置信息,实现精确定位;其次在频域引入细粒度频率注意力模块,提取包含纹理、边缘等全局细微伪造信息的高频特征。尽管两个域的特征可独立优化,但直接融合效果有限。为此,我们设计特征叠加策略,将特征转为波形令牌并基于相位更新,有效放大真实与伪造特征的差异。该方法在五个公开深度伪造数据集上显著优于当前最优(SOTA)方法,在多种篡改操作及真实场景下均表现出更强的异常捕获能力。
原文摘要 · Abstract (English)
Deepfake detection is crucial for curbing the harm it causes to society. However, current Deepfake detection methods fail to thoroughly explore artifact information across different domains due to insufficient intrinsic interactions. These interactions refer to the fusion and coordination after feature extraction processes across different domains, which are crucial for recognizing complex forgery clues. Focusing on more generalized Deepfake detection, in this work, we introduce a novel bi-directional attention module to capture the local positional information of artifact clues from the spatial domain. This enables accurate artifact localization, thus addressing the coarse processing with artifact features. To further address the limitation that the proposed bi-directional attention module may not well capture global subtle forgery information in the artifact feature (e.g., textures or edges), we employ a fine-grained frequency attention module in the frequency domain. By doing so, we can obtain high-frequency information in the fine-grained features, which contains the global and subtle forgery information. Although these features from the diverse domains can be effectively and independently improved, fusing them directly does not effectively improve the detection performance. Therefore, we propose a feature superposition strategy that complements information from spatial and frequency domains. This strategy turns the feature components into the form of wave-like tokens, which are updated based on their phase, such that the distinctions between authentic and artifact features can be amplified. Our method demonstrates significant improvements over state-of-the-art (SOTA) methods on five public Deepfake datasets in capturing abnormalities across different manipulated operations and real-life.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。