arXiv:2605.24965cs.CVcs.AI2026-05

探究视觉大模型在跨域伪造人脸检测中的泛化极限

Cross-Domain Generalization Limits of Vision Foundation Models in Facial Deepfake Detection

  • 对比三种大模型在跨域检测中的表现
  • 发现模型对局部编辑类伪造泛化能力有限
  • 适合关注深度伪造检测泛化性的研究者

生成模型的快速发展催生了高度逼真的面部深度伪造,暴露出现代数字取证的关键弱点:检测器难以泛化到未见的篡改技术。传统网络存在表征坍塌问题,过度依赖特定训练生成器的局部伪影特征。本文系统评估现代视觉基础模型作为即用型特征提取器,在跨生成流形追踪取证异常方面的通用性。通过在挑战性DF40基准上部署冻结主干进行下游线性探测,对比了三种基础学习范式:全监督宏观语义特征(RoPE-ViT)、纯自监督几何特征(DINOv3)以及多教师聚合表示(NVIDIA C-RADIOv4-H)。实证结果揭示预训练范式与参数规模间的内在权衡,证明尽管基础模型在整体人脸合成任务中保持高判别力,但在局部人脸编辑技术面前暴露出线性探测结构的根本局限。

原文摘要 · Abstract (English)

The rapid evolution of generative models has enabled the creation of hyper-realistic facial deepfakes, exposing a critical vulnerability in modern digital forensics: the inability of detectors to generalize to unseen manipulation techniques. Traditional networks suffer from representation collapse, overfitting to localized artifact fingerprints of specific training generators. This work investigates whether modern Vision Foundation Models can serve as generalizable, out-of-the-box feature extractors capable of tracking forensic anomalies across entirely unseen generative manifolds. We conduct a systematic cross-domain evaluation comparing three foundational learning paradigms: fully supervised macro-semantic features (RoPE-ViT), pure self-supervised geometric features (DINOv3), and multi-teacher agglomerative representations (NVIDIA C-RADIOv4-H). By deploying frozen backbones subjected to downstream linear probing, we map the performance limitations of these architectures on the challenging DF40 benchmark. Our empirical findings expose the intrinsic trade-offs between pre-training paradigms and parameter scale, proving that while foundation models retain high discriminative capabilities for entire face synthesis, localized face editing techniques expose fundamental boundaries in linear probe evaluation structures. Source code and model weights are available in http://github.com/mribrahim/deepfake

深度伪造视觉模型泛化性检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。