arXiv:2505.02013cs.CV2025-05被引 17

用视觉语言融合提升深度伪造人脸检测准确率

MLLM-Enhanced Face Forgery Detection: A Vision-Language Fusion Solution

  • 设计双向交互的视觉语言网络,融合图像与文本特征
  • 在跨数据集和单数据集测试中均达当前最优性能
  • 新增带解释性标注的FF++扩展数据集,助力模型训练

可靠的面部伪造检测算法对于应对深度伪造引发的虚假信息威胁至关重要。以往研究显示多模态大语言模型(MLLM)在识别篡改人脸方面具有潜力,但现有方法通常依赖大语言模型(LLM)或外部检测器独立生成分类结果,导致视觉与文本模态融合不充分。本文提出VLF-FFD,一种新型的视觉-语言融合解决方案。主要贡献有二:其一,构建了EFF++,即对广泛使用的FaceForensics++(FF++)数据集的帧级扩展,每帧篡改视频均配有描述伪造痕迹及具体操作技术的文本注释,支持更有效、更具解释性的MLLM训练;其二,设计了视觉-语言融合网络(VLF-Net),通过三阶段训练流程促进视觉与文本特征的双向交互。VLF-FFD在跨数据集与单数据集评估中均达到当前最优表现,证明其在人脸伪造检测中的卓越有效性。

原文摘要 · Abstract (English)

Reliable face forgery detection algorithms are crucial for countering the growing threat of deepfake-driven disinformation. Previous research has demonstrated the potential of Multimodal Large Language Models (MLLMs) in identifying manipulated faces. However, existing methods typically depend on either the Large Language Model (LLM) alone or an external detector to generate classification results, which often leads to sub-optimal integration of visual and textual modalities. In this paper, we propose VLF-FFD, a novel Vision-Language Fusion solution for MLLM-enhanced Face Forgery Detection. Our key contributions are twofold. First, we present EFF++, a frame-level, explainability-driven extension of the widely used FaceForensics++ (FF++) dataset. In EFF++, each manipulated video frame is paired with a textual annotation that describes both the forgery artifacts and the specific manipulation technique applied, enabling more effective and informative MLLM training. Second, we design a Vision-Language Fusion Network (VLF-Net) that promotes bidirectional interaction between visual and textual features, supported by a three-stage training pipeline to fully leverage its potential. VLF-FFD achieves state-of-the-art (SOTA) performance in both cross-dataset and intra-dataset evaluations, underscoring its exceptional effectiveness in face forgery detection.

伪造检测视觉语言多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。