arXiv:2603.29295cs.CV2026-03

用眼神特征提升假脸检测,让模型更懂新生成手法。

GazeCLIP: Gaze-Guided CLIP with Adaptive-Enhanced Fine-Grained Language Prompt for Deepfake Attribution and Detection

  • 基于眼神差异设计视觉编码器,融合外观与眼神信息
  • 在新生成模型上准确率提升6.56%,AUC提高5.32%
  • 适合关注假脸溯源与检测泛化能力的研究者

现有深度伪造溯源与检测方法在面对新型生成技术时泛化能力差,仅依赖视觉模态且粗略评估未见生成器的性能,未能考虑任务间的协同。为此,我们提出一种基于眼神引导的CLIP模型,结合自适应增强细粒度语言提示,实现细粒度深度伪造溯源与检测(DFAD)。我们构建了新的精细基准,评估网络在扩散模型、流模型等新型生成器上的表现。基于真实与伪造眼神向量分布差异显著、以及生成图像中目标眼神保留程度不同的观察,设计视觉感知编码器,利用眼神差异挖掘跨外观与眼神域的全局伪造嵌入。提出眼神感知图像编码器(GIE),融合眼神提示与通用伪造图像嵌入,使特征进入更稳定统一的DFAD特征空间。构建语言精炼编码器(LRE),通过自适应增强词选择器生成动态优化的语言嵌入,实现精准视觉-语言匹配。大量实验表明,在新生成器上,本模型在溯源与检测任务中平均准确率提升6.56%,AUC提升5.32%。代码将开源于GitHub。

原文摘要 · Abstract (English)

Current deepfake attribution or deepfake detection works tend to exhibit poor generalization to novel generative methods due to the limited exploration in visual modalities alone. They tend to assess the attribution or detection performance of models on unseen advanced generators, coarsely, and fail to consider the synergy of the two tasks. To this end, we propose a novel gaze-guided CLIP with adaptive-enhanced fine-grained language prompts for fine-grained deepfake attribution and detection (DFAD). Specifically, we conduct a novel and fine-grained benchmark to evaluate the DFAD performance of networks on novel generators like diffusion and flow models. Additionally, we introduce a gaze-aware model based on CLIP, which is devised to enhance the generalization to unseen face forgery attacks. Built upon the novel observation that there are significant distribution differences between pristine and forged gaze vectors, and the preservation of the target gaze in facial images generated by GAN and diffusion varies significantly, we design a visual perception encoder to employ the inherent gaze differences to mine global forgery embeddings across appearance and gaze domains. We propose a gaze-aware image encoder (GIE) that fuses forgery gaze prompts extracted via a gaze encoder with common forged image embeddings to capture general attribution patterns, allowing features to be transformed into a more stable and common DFAD feature space. We build a language refinement encoder (LRE) to generate dynamically enhanced language embeddings via an adaptive-enhanced word selector for precise vision-language matching. Extensive experiments on our benchmark show that our model outperforms the state-of-the-art by 6.56% ACC and 5.32% AUC in average performance under the attribution and detection settings, respectively. Codes will be available on GitHub.

假脸检测视觉语言模型眼神分析泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。