arXiv:2503.20188cs.CV2025-03CVPR被引 59

同时检测伪造人脸并生成可解释文本,提升识别准确率与透明度

Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face Detector

论文配图:Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face Detector
图 1 · 摘自论文原文
  • 结合CLIP与大语言模型,实现多模态融合检测与解释
  • 在多种伪造数据集上达到顶尖检测性能,支持未见过的伪造类型
  • 适合需要可解释性的AI安全场景,如社交媒体审核

深度伪造检测是防范恶意虚假信息传播的关键研究方向。现有方法通常仅提供二分类结果或独立的文本解释,无法兼顾准确性与可解释性。本文提出一种新型多模态伪造人脸检测器M2F2-Det,利用预训练的CLIP模型增强对未知伪造类型的泛化能力,并引入大语言模型(LLM)生成详细文本解释,使检测决策更具可理解性。该方法通过定制化伪造人脸提示学习策略,有效融合视觉与语言模态信息。实验表明,M2F2-Det在检测任务和解释生成任务中均达到当前最优性能,能够准确识别多种伪造手法并生成自然语言解释。

原文摘要 · Abstract (English)

Deepfake detection is a long-established research topic vital for mitigating the spread of malicious misinformation. Unlike prior methods that provide either binary classification results or textual explanations separately, we introduce a novel method capable of generating both simultaneously. Our method harnesses the multi-modal learning capability of the pre-trained CLIP and the unprecedented interpretability of large language models (LLMs) to enhance both the generalization and explainability of deepfake detection. Specifically, we introduce a multi-modal face forgery detector (M2F2-Det) that employs tailored face forgery prompt learning, incorporating the pre-trained CLIP to improve generalization to unseen forgeries. Also, M2F2-Det incorporates an LLM to provide detailed textual explanations of its detection decisions, enhancing interpretability by bridging the gap between natural language and subtle cues of facial forgeries. Empirically, we evaluate M2F2-Det on both detection and explanation generation tasks, where it achieves state-of-the-art performance, demonstrating its effectiveness in identifying and explaining diverse forgeries.

伪造检测多模态可解释性CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。