arXiv:2508.05526cs.CV2025-08中稿 · KDD被引 2

用图神经网络统一检测视频伪造,轻量高效。

When Deepfake Detection Meets Graph Neural Network:a Unified and Lightweight Learning Framework

  • 将视频建模为图结构,联合分析空间、时间、频域异常
  • 参数量比顶尖模型少42倍,跨域检测性能更优
  • 适合边缘设备部署,兼顾精度与计算效率

生成式视频模型的泛滥使检测AI生成或篡改视频成为紧迫挑战。现有方法多依赖孤立的空间、时间或频域信息,难以跨类型泛化,且通常需要大型模型。本文提出SSTGNN——一种轻量级空间-频域-时序图神经网络框架,将视频表示为结构化图,实现对空间不一致、时间伪影和频域失真的联合推理。SSTGNN融合可学习频域滤波器与时空差分建模,更有效捕捉细微篡改痕迹。在多个基准数据集上的实验表明,SSTGNN在同域与跨域设置下均表现优异,同时具备强效率与资源友好性。令人瞩目的是,其参数量相比当前最优模型减少高达42倍,适用于真实场景的轻量级部署。

原文摘要 · Abstract (English)

The proliferation of generative video models has made detecting AI-generated and manipulated videos an urgent challenge. Existing detection approaches often fail to generalize across diverse manipulation types due to their reliance on isolated spatial, temporal, or spectral information, and typically require large models to perform well. This paper introduces SSTGNN, a lightweight Spatial-Spectral-Temporal Graph Neural Network framework that represents videos as structured graphs, enabling joint reasoning over spatial inconsistencies, temporal artifacts, and spectral distortions. SSTGNN incorporates learnable spectral filters and spatial-temporal differential modeling into a unified graph-based architecture, capturing subtle manipulation traces more effectively. Extensive experiments on diverse benchmark datasets demonstrate that SSTGNN not only achieves superior performance in both in-domain and cross-domain settings, but also offers strong efficiency and resource allocation. Remarkably, SSTGNN accomplishes these results with up to 42$\times$ fewer parameters than state-of-the-art models, making it highly lightweight and resource-friendly for real-world deployment.

深度伪造检测图神经网络轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。