arXiv:2501.01184cs.CV2025-01ICCV被引 14

通过细粒度时空分析,提升深伪视频检测的泛化能力。

Vulnerability-Aware Spatio-Temporal Learning for Generalizable Deepfake Video Detection

  • 引入多任务学习框架,聚焦易出错的时空区域
  • 生成带细微伪造痕迹的伪假视频,提供高质量训练样本
  • 适合需要强泛化能力的深伪检测场景

深伪视频检测面临复杂时空伪影建模的挑战。现有方法多依赖真实与虚假图像序列的二分类训练,难以泛化到未见生成方法。随着生成AI持续进步,深伪痕迹在空间和时间层面日益难以察觉。为此,我们提出名为FakeSTormer的细粒度检测方法,通过建模微弱时空不一致来避免过拟合。具体包括:设计多任务学习框架,引入两个辅助分支以显式关注易产生伪影的空间与时间区域;提出视频级数据合成策略,生成带有细微时空伪影的伪假视频,为辅助分支提供高质量样本及无需人工标注的数据。在多个挑战性基准上的大量实验表明,该方法优于近期最先进方法。代码已公开于https://github.com/10Ring/FakeSTormer。

原文摘要 · Abstract (English)

Detecting deepfake videos is highly challenging given the complexity of characterizing spatio-temporal artifacts. Most existing methods rely on binary classifiers trained using real and fake image sequences, therefore hindering their generalization capabilities to unseen generation methods. Moreover, with the constant progress in generative Artificial Intelligence (AI), deepfake artifacts are becoming imperceptible at both the spatial and the temporal levels, making them extremely difficult to capture. To address these issues, we propose a fine-grained deepfake video detection approach called FakeSTormer that enforces the modeling of subtle spatio-temporal inconsistencies while avoiding overfitting. Specifically, we introduce a multi-task learning framework that incorporates two auxiliary branches for explicitly attending artifact-prone spatial and temporal regions. Additionally, we propose a video-level data synthesis strategy that generates pseudo-fake videos with subtle spatio-temporal artifacts, providing high-quality samples and hand-free annotations for our additional branches. Extensive experiments on several challenging benchmarks demonstrate the superiority of our approach compared to recent state-of-the-art methods. The code is available at https://github.com/10Ring/FakeSTormer.

深伪检测时空建模泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。