arXiv:2507.18031cs.CVcs.AI2025-07被引 1

用视觉语言模型解释+图神经网络,提升深伪图像检测能力

ViGText: Deepfake Image Detection with Vision-Language Model Explanations and Graph Neural Networks

  • 将图像分块构建图文图,用GNN融合分析
  • 通用性测试F1达98.32%,定制化深伪检测显著提升
  • 对针对性攻击鲁棒性强,性能下降不足4%

深度伪造技术的快速发展对媒体真实性构成威胁。传统检测方法在面对复杂、定制化的深伪图像时,普遍存在泛化能力和抗攻击鲁棒性不足的问题。本文提出ViGText,通过在图神经网络框架中融合视觉大语言模型(VLLM)生成的详细文本解释与图像数据,实现更精准的深伪检测。该方法将图像划分为图像块,构建图像与文本图,并利用多尺度特征提取(空间与频域)捕捉细微不一致。实验表明,ViGText在泛化评估中平均F1分数从72.45%提升至98.32%,展现出对未见微调扩散模型生成内容的强大适应能力;在鲁棒性方面,召回率相比其他方法提升11.1%;面对针对图结构的定向攻击,性能下降低于4%。该方法通过深入的图文联合分析,为保障媒体真实性和信息完整性树立新标准。

原文摘要 · Abstract (English)

The rapid rise of deepfake technology, which produces realistic but fraudulent digital content, threatens the authenticity of media. Traditional deepfake detection approaches often struggle with sophisticated, customized deepfakes, especially in terms of generalization and robustness against malicious attacks. This paper introduces ViGText, a novel approach that integrates images with Vision Large Language Model (VLLM) Text explanations within a Graph-based framework to improve deepfake detection. The novelty of ViGText lies in its integration of detailed explanations with visual data, as it provides a more context-aware analysis than captions, which often lack specificity and fail to reveal subtle inconsistencies. ViGText systematically divides images into patches, constructs image and text graphs, and integrates them for analysis using Graph Neural Networks (GNNs) to identify deepfakes. Through the use of multi-level feature extraction across spatial and frequency domains, ViGText captures details that enhance its robustness and accuracy to detect sophisticated deepfakes. Extensive experiments demonstrate that ViGText significantly enhances generalization and achieves a notable performance boost when it detects user-customized deepfakes. Specifically, average F1 scores rise from 72.45% to 98.32% under generalization evaluation, and reflects the model's superior ability to generalize to unseen, fine-tuned variations of stable diffusion models. As for robustness, ViGText achieves an increase of 11.1% in recall compared to other deepfake detection approaches. When facing targeted attacks that exploit its graph-based architecture, ViGText limits classification performance degradation to less than 4%. ViGText uses detailed visual and textual analysis to set a new standard for detecting deepfakes, helping ensure media authenticity and information integrity.

深伪检测图文模型图神经网络鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。