arXiv:2602.15903cs.CV2026-02

用多方法融合与图文对齐提升深伪检测泛化能力

Detecting Deepfakes with Multivariate Soft Blending and CLIP-based Image-Text Alignment

  • 通过随机混合多种伪造方法生成训练样本,增强模型泛化性
  • 在跨数据集测试中平均AUC提升3.27%,优于现有方法
  • 适合需要高鲁棒性检测的安防与内容审核场景

深度伪造技术的泛滥亟需可靠的检测方法。然而,现有方法常因不同伪造手段间分布差异大而出现准确率低、泛化性差的问题。为此,我们提出一种基于CLIP引导伪造强度估计的多变量软混合增强框架(MSBA-CLIP)。该方法利用CLIP的多模态对齐能力捕捉细微伪造痕迹,设计多变量软混合增强策略,通过随机加权融合多种伪造图像,迫使模型学习通用特征。同时引入多变量伪造强度估计模块,显式指导模型学习不同伪造模式与强度下的特征。大量实验表明,该方法在域内测试中准确率和AUC分别较最优基线提升3.32%和4.02%;在五个跨域数据集上平均AUC提升3.27%。消融实验验证了各组件有效性。尽管依赖大型视觉语言模型导致计算成本较高,但本工作显著推动了更通用、更鲁棒的深伪检测发展。

原文摘要 · Abstract (English)

The proliferation of highly realistic facial forgeries necessitates robust detection methods. However, existing approaches often suffer from limited accuracy and poor generalization due to significant distribution shifts among samples generated by diverse forgery techniques. To address these challenges, we propose a novel Multivariate and Soft Blending Augmentation with CLIP-guided Forgery Intensity Estimation (MSBA-CLIP) framework. Our method leverages the multimodal alignment capabilities of CLIP to capture subtle forgery traces. We introduce a Multivariate and Soft Blending Augmentation (MSBA) strategy that synthesizes images by blending forgeries from multiple methods with random weights, forcing the model to learn generalizable patterns. Furthermore, a dedicated Multivariate Forgery Intensity Estimation (MFIE) module is designed to explicitly guide the model in learning features related to varied forgery modes and intensities. Extensive experiments demonstrate state-of-the-art performance. On in-domain tests, our method improves Accuracy and AUC by 3.32\% and 4.02\%, respectively, over the best baseline. In cross-domain evaluations across five datasets, it achieves an average AUC gain of 3.27\%. Ablation studies confirm the efficacy of both proposed components. While the reliance on a large vision-language model entails higher computational cost, our work presents a significant step towards more generalizable and robust deepfake detection.

深伪检测图文对齐多模态泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。