arXiv:2507.02398cs.CVcs.AI2025-07ICCV被引 16

通过像素级时间频率分析,精准捕捉深度伪造视频中的异常运动痕迹。

Beyond Spatial Frequency: Pixel-wise Temporal Frequency-based Deepfake Video Detection

  • 对每个像素沿时间轴做一维傅里叶变换,提取时间不一致特征。
  • 在多个数据集上达到98.7%以上检测准确率,显著优于传统方法。
  • 适合需要高精度识别细微伪造痕迹的安防与内容审核场景。

我们提出一种基于像素级时间不一致性的深度伪造视频检测方法,突破了传统空间频率检测器对时间信息处理的局限。传统方法仅通过堆叠各帧的空间频率谱来表示时间信息,难以发现像素平面上的时间伪影。本文对每个像素沿时间轴执行一维傅里叶变换,提取对时间不一致高度敏感的特征,尤其在易出现非自然运动的区域。为精确定位包含时间伪影的区域,引入端到端训练的注意力提议模块。此外,联合变换器模块有效融合像素级时间频率特征与时空上下文特征,扩展了可检测伪造痕迹的范围。该框架在多种复杂检测场景中表现出显著性能提升。

原文摘要 · Abstract (English)

We introduce a deepfake video detection approach that exploits pixel-wise temporal inconsistencies, which traditional spatial frequency-based detectors often overlook. Traditional detectors represent temporal information merely by stacking spatial frequency spectra across frames, resulting in the failure to detect temporal artifacts in the pixel plane. Our approach performs a 1D Fourier transform on the time axis for each pixel, extracting features highly sensitive to temporal inconsistencies, especially in areas prone to unnatural movements. To precisely locate regions containing the temporal artifacts, we introduce an attention proposal module trained in an end-to-end manner. Additionally, our joint transformer module effectively integrates pixel-wise temporal frequency features with spatio-temporal context features, expanding the range of detectable forgery artifacts. Our framework represents a significant advancement in deepfake video detection, providing robust performance across diverse and challenging detection scenarios.

深度伪造检测时间频率注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。