arXiv:2503.02503cs.CV2025-03被引 2

通过知识注入提升深度伪造检测的泛化能力

Deepfake Detection via Knowledge Injection

  • 构建多任务知识注入框架,融合真实数据知识
  • 在多个ViT模型上实现最佳泛化性能,收敛更快
  • 适合需要强泛化能力的深度伪造检测场景

深度伪造检测技术因生成式AI可制造逼真伪造内容而愈发重要。现有方法或依赖分类模型拟合训练数据分布,或利用伪造生成机制学习伪造分布,但普遍忽视真实数据知识,导致对未见数据的泛化能力不足。为此,本文提出基于知识注入的深度伪造检测方法(KID),通过多任务学习框架将必要知识注入现有ViT类主干模型,包括基础模型。设计知识注入模块以更准确建模真实与伪造数据分布,并引入粗粒度伪造定位分支,在多任务学习中丰富伪造知识。提出两阶段层间抑制与对比损失,强化真实数据知识,平衡真实与伪造知识比例。大量实验表明,KID在不同规模ViT主干模型上具备良好兼容性,实现领先泛化性能并加快训练收敛速度。

原文摘要 · Abstract (English)

Deepfake detection technologies become vital because current generative AI models can generate realistic deepfakes, which may be utilized in malicious purposes. Existing deepfake detection methods either rely on developing classification methods to better fit the distributions of the training data, or exploiting forgery synthesis mechanisms to learn a more comprehensive forgery distribution. Unfortunately, these methods tend to overlook the essential role of real data knowledge, which limits their generalization ability in processing the unseen real and fake data. To tackle these challenges, in this paper, we propose a simple and novel approach, named Knowledge Injection based deepfake Detection (KID), by constructing a multi-task learning based knowledge injection framework, which can be easily plugged into existing ViT-based backbone models, including foundation models. Specifically, a knowledge injection module is proposed to learn and inject necessary knowledge into the backbone model, to achieve a more accurate modeling of the distributions of real and fake data. A coarse-grained forgery localization branch is constructed to learn the forgery locations in a multi-task learning manner, to enrich the learned forgery knowledge for the knowledge injection module. Two layer-wise suppression and contrast losses are proposed to emphasize the knowledge of real data in the knowledge injection module, to further balance the portions of the real and fake knowledge. Extensive experiments have demonstrated that our KID possesses excellent compatibility with different scales of Vit-based backbone models, and achieves state-of-the-art generalization performance while enhancing the training convergence speed.

深度伪造知识注入ViT检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。