通过空间频率协同分析,提升伪造视频检测的泛化能力。
Towards Generalizable Deepfake Detection with Spatial-Frequency Collaborative Learning and Hierarchical Cross-Modal Fusion
- 融合多尺度空间与频域特征,捕捉细微伪造痕迹。
- 在多个基准上准确率优于现有方法,泛化性能显著提升。
- 适合需要应对新型伪造手段的安全检测场景。
深度生成模型的快速演进给深度伪造检测带来严峻挑战,现有检测器因依赖特定伪造痕迹,在面对未见伪造类型时性能急剧下降。尽管已有方法主要关注空间域分析,频域操作仍局限于特征层面增强,导致频域原生特征及空间-频域交互作用未被充分挖掘。为此,我们提出一种新型检测框架,集成多尺度空间-频域分析以实现通用深度伪造检测。该框架包含三个核心组件:(1) 局部谱特征提取管道,结合分块离散余弦变换与级联多尺度卷积,捕捉细微谱特征;(2) 全局谱特征提取管道,利用尺度不变差分累积识别整体伪造分布模式;(3) 多阶段跨模态融合机制,结合浅层注意力增强与深层动态调制,建模空间-频域交互关系。在广泛采用的基准上进行的大量评估表明,该方法在准确率和泛化能力方面均优于当前最优深度伪造检测方法。
原文摘要 · Abstract (English)
The rapid evolution of deep generative models poses a critical challenge to deepfake detection, as detectors trained on forgery-specific artifacts often suffer significant performance degradation when encountering unseen forgeries. While existing methods predominantly rely on spatial domain analysis, frequency domain operations are primarily limited to feature-level augmentation, leaving frequency-native artifacts and spatial-frequency interactions insufficiently exploited. To address this limitation, we propose a novel detection framework that integrates multi-scale spatial-frequency analysis for universal deepfake detection. Our framework comprises three key components: (1) a local spectral feature extraction pipeline that combines block-wise discrete cosine transform with cascaded multi-scale convolutions to capture subtle spectral artifacts; (2) a global spectral feature extraction pipeline utilizing scale-invariant differential accumulation to identify holistic forgery distribution patterns; and (3) a multi-stage cross-modal fusion mechanism that incorporates shallow-layer attention enhancement and deep-layer dynamic modulation to model spatial-frequency interactions. Extensive evaluations on widely adopted benchmarks demonstrate that our method outperforms state-of-the-art deepfake detection methods in both accuracy and generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。