arXiv:2605.07398cs.CVcs.AI2026-05

提出防御深度伪造视频时间攻击的新方法,提升检测鲁棒性。

Exposing and Mitigating Temporal Attack in Deepfake Video Detection

论文配图:Exposing and Mitigating Temporal Attack in Deepfake Video Detection
图 1 · 摘自论文原文
  • 设计可学习的时频对抗样本,模拟极端攻击场景
  • 在模拟幅度谱攻击下AUC提升21.30个百分点
  • 适合关注视频伪造检测安全性的研究者

尽管时空深度伪造检测器已实现高AUC,但实验显示其易受规避攻击。这些模型倾向于过度依赖脆弱的时间频谱特征,而非学习稳健的语义因果关系。为此,我们提出SpInShield,一种显式解耦语义运动与可操纵频谱伪影的时频不变防御框架。设计可学习的频谱对抗器,动态合成严重频谱畸变以模拟极端攻击场景。通过快捷路径抑制优化策略,促使编码器提取可靠取证特征,并清除潜在空间中的不稳频谱统计量。实验表明,SpInShield在常用数据集上表现优异,在模拟幅度谱攻击下相比最强基线提升21.30个百分点AUC。

原文摘要 · Abstract (English)

While spatiotemporal deepfake detectors achieve high AUC, our experiments reveal their susceptibility to evasion attacks. These models tend to overfit on fragile temporal spectrum cues, rather than learning robust semantic causality. To mitigate this vulnerability, we propose SpInShield, a temporal spectral-invariant defense framework explicitly designed to decouple semantic motion from manipulatable spectral artifacts. We propose a learnable spectral adversary that dynamically synthesizes severe spectral deformations, simulating extreme attack scenarios. By employing a shortcut suppression optimization strategy, SpInShield compels the encoder to extract reliable forensic cues while purging unstable spectral statistics from the latent space. Experiments show that SpInShield obtains competitive performance on widely used datasets and outperforms the strongest baseline by 21.30 percentage points in AUC under simulated amplitude spectral attacks.

深度伪造视频检测对抗防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。