通过消除空间依赖,提升视频级深度伪造检测的泛化能力
Reduced Spatial Dependency for More General Video-level Deepfake Detection
- 设计多空间扰动分支,构建受干扰的特征簇
- 利用互信息原理融合跨簇时序特征,提取内在时间规律
- 适合需要强泛化能力的视频安全检测场景
作为主流的AI生成内容之一,深度伪造引发了显著的安全隐患。尽管已有研究证明时序一致性线索具有更好的泛化能力,但基于CNN的方法不可避免地引入空间偏差,阻碍了对内在时序特征的提取。为此,我们提出一种名为空间依赖性降低(SDR)的新方法,通过整合多个空间扰动簇中的共同时序一致性特征,减少模型对空间信息的依赖。具体而言,设计多个空间扰动分支(SPB)以构建空间扰动的特征簇;随后,基于互信息理论,提出任务相关特征融合(TRFI)模块,从这些簇中捕捉位于相似潜在空间的时序特征;最后,将融合后的特征输入时序变压器以捕获长程依赖关系。大量基准测试与消融实验验证了该方法的有效性与合理性。
原文摘要 · Abstract (English)
As one of the prominent AI-generated content, Deepfake has raised significant safety concerns. Although it has been demonstrated that temporal consistency cues offer better generalization capability, existing methods based on CNNs inevitably introduce spatial bias, which hinders the extraction of intrinsic temporal features. To address this issue, we propose a novel method called Spatial Dependency Reduction (SDR), which integrates common temporal consistency features from multiple spatially-perturbed clusters, to reduce the dependency of the model on spatial information. Specifically, we design multiple Spatial Perturbation Branch (SPB) to construct spatially-perturbed feature clusters. Subsequently, we utilize the theory of mutual information and propose a Task-Relevant Feature Integration (TRFI) module to capture temporal features residing in similar latent space from these clusters. Finally, the integrated feature is fed into a temporal transformer to capture long-range dependencies. Extensive benchmarks and ablation studies demonstrate the effectiveness and rationale of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。