arXiv:2410.06126cs.CV2024-10NeurIPS被引 39

用多模态大模型提升伪造视频检测的可解释性与扩展性

X2-DFD: A framework for eXplainable and eXtendable Deepfake Detection

论文配图:X2-DFD: A framework for eXplainable and eXtendable Deepfake Detection
图 1 · 摘自论文原文
  • 基于多模态大模型构建三阶段框架,评估并强化检测特征
  • 在多个数据集上显著提升检测准确率,可解释性优于原始模型
  • 模块化设计支持未来升级,适合安全与内容审核场景

本文提出X2-DFD,一种基于多模态大语言模型(MLLM)的可解释且可扩展的深度伪造检测框架,包含三个关键阶段。第一阶段模型特征评估系统性地衡量伪造相关特征对MLLM的可检测性,生成基于特征重要性的优先级排序。第二阶段可解释数据集构建包含两个模块:强特征强化通过增强模型已学习特征来提升检测与解释能力;弱特征补充则通过引入特定特征探测器(如低层伪影分析器)弥补MLLM的不足。第三阶段为微调与推理,在构建的数据集上微调MLLM并部署进行最终检测与解释。该方法融合三阶段优势,既强化了模型长处又补足短板,显著提升可检测性与可解释性。大量实验与消融分析,结合全面的人类评估,验证了该方法相比原始MLLM的优越性能。更令人鼓舞的是,该框架具备即插即用特性,可无缝集成未来更先进的MLLM和特定特征探测器,实现持续演进以应对快速演变的深度伪造挑战。

原文摘要 · Abstract (English)

This paper proposes X2-DFD, an eXplainable and eXtendable framework based on multimodal large-language models (MLLMs) for deepfake detection, consisting of three key stages. The first stage, Model Feature Assessment, systematically evaluates the detectability of forgery-related features for the MLLM, generating a prioritized ranking of features based on their intrinsic importance to the model. The second stage, Explainable Dataset Construction, consists of two key modules: Strong Feature Strengthening, which is designed to enhance the model's existing detection and explanation capabilities by reinforcing its well-learned features, and Weak Feature Supplementing, which addresses gaps by integrating specific feature detectors (e.g., low-level artifact analyzers) to compensate for the MLLM's limitations. The third stage, Fine-tuning and Inference, involves fine-tuning the MLLM on the constructed dataset and deploying it for final detection and explanation. By integrating these three stages, our approach enhances the MLLM's strengths while supplementing its weaknesses, ultimately improving both the detectability and explainability. Extensive experiments and ablations, followed by a comprehensive human study, validate the improved performance of our approach compared to the original MLLMs. More encouragingly, our framework is designed to be plug-and-play, allowing it to seamlessly integrate with future more advanced MLLMs and specific feature detectors, leading to continual improvement and extension to face the challenges of rapidly evolving deepfakes.

深度伪造可解释性多模态检测框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。