arXiv:2508.19730cs.CV2025-08

用人脸基础模型提升深度伪造检测泛化能力

Improving Generalization in Deepfake Detection with Face Foundation Models and Metric Learning

  • 基于自监督人脸模型FSFM,融合多种伪造数据微调
  • 引入三元组损失,增强真实与伪造样本的特征区分度
  • 支持按篡改类型分类监督,适合真实场景下的检测需求

随着深度伪造技术日益逼真且易于获取,媒体真实性与信息完整性面临严峻挑战。尽管已有进展,现有检测模型在分布外场景下泛化能力仍不足。本文提出一种强泛化视频深度伪造检测框架,利用在真实人脸数据上训练的自监督模型FSFM学习丰富的人脸表征,并在其基础上,使用涵盖人脸替换与重演的多数据集组合进行微调。为增强判别力,训练中引入三元组损失变体,引导模型生成更可分的真实与伪造嵌入。此外,探索基于归因的监督策略,将伪造内容按篡改类型或来源数据集分类,评估其对泛化性能的影响。大量实验表明,该方法在多样化的评估基准上表现优异,尤其在复杂真实场景中具备显著优势。

原文摘要 · Abstract (English)

The increasing realism and accessibility of deepfakes have raised critical concerns about media authenticity and information integrity. Despite recent advances, deepfake detection models often struggle to generalize beyond their training distributions, particularly when applied to media content found in the wild. In this work, we present a robust video deepfake detection framework with strong generalization that takes advantage of the rich facial representations learned by face foundation models. Our method is built on top of FSFM, a self-supervised model trained on real face data, and is further fine-tuned using an ensemble of deepfake datasets spanning both face-swapping and face-reenactment manipulations. To enhance discriminative power, we incorporate triplet loss variants during training, guiding the model to produce more separable embeddings between real and fake samples. Additionally, we explore attribution-based supervision schemes, where deepfakes are categorized by manipulation type or source dataset, to assess their impact on generalization. Extensive experiments across diverse evaluation benchmarks demonstrate the effectiveness of our approach, especially in challenging real-world scenarios.

深度伪造检测人脸基础模型度量学习泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。