arXiv:2510.16442cs.CVcs.AI2025-10被引 4

用大模型推理实现可解释的深度伪造视频检测。

EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning

  • 通过时空细粒度特征提取,融合跨帧伪造线索。
  • 引入人脸特征约束,实现像素级定位与可靠推理。
  • 支持跨伪造方法和跨数据集检测,适合可信安全场景。

深度伪造视频技术的快速发展既推动了艺术创作,也加剧了虚假信息传播。传统检测方法存在原理不透明、泛化能力弱等问题。本文提出可解释深度伪造视频检测(EDVD)任务,设计基于多模态大语言模型(MLLM)的EDVD-LLaMA框架,实现高精度检测与可追溯的推理解释。首先,采用时空细微信息分词(ST-SIT)提取并融合全局与局部跨帧伪造特征,为模型提供丰富的时空语义输入。其次,构建细粒度多模态思维链(Fg-MCoT),在推理中引入人脸特征作为硬约束,实现像素级时空定位,抑制幻觉输出,提升推理可靠性。此外,构建可解释推理FF++数据集(ER-FF++set),通过结构化标注确保数据质量,支持推理与检测双重监督。大量实验表明,EDVD-LLaMA在检测精度、可解释性及跨伪造方法、跨数据集场景下均表现优异,显著优于现有方法。项目主页:https://11ouo1.github.io/edvd-llama/

原文摘要 · Abstract (English)

The rapid development of deepfake video technology has not only facilitated artistic creation but also made it easier to spread misinformation. Traditional deepfake video detection (DVD) methods face issues such as a lack of transparency in their principles and insufficient generalization capabilities to cope with evolving forgery techniques. This highlights an urgent need for detectors that can identify forged content and provide verifiable reasoning explanations. This paper proposes the explainable deepfake video detection (EDVD) task and designs the EDVD-LLaMA multimodal, a large language model (MLLM) reasoning framework, which provides traceable reasoning processes alongside accurate detection results and trustworthy explanations. Our approach first incorporates a Spatio-Temporal Subtle Information Tokenization (ST-SIT) to extract and fuse global and local cross-frame deepfake features, providing rich spatio-temporal semantic information input for MLLM reasoning. Second, we construct a Fine-grained Multimodal Chain-of-Thought (Fg-MCoT) mechanism, which introduces facial feature data as hard constraints during the reasoning process to achieve pixel-level spatio-temporal video localization, suppress hallucinated outputs, and enhance the reliability of the chain of thought. In addition, we build an Explainable Reasoning FF++ dataset (ER-FF++set), leveraging structured data to annotate videos and ensure quality control, thereby supporting dual supervision for reasoning and detection. Extensive experiments demonstrate that EDVD-LLaMA achieves outstanding performance and robustness in terms of detection accuracy, explainability, and its ability to handle cross-forgery methods and cross-dataset scenarios. Compared to previous DVD methods, it provides a more explainable and superior solution. The project page is available at: https://11ouo1.github.io/edvd-llama/.

深度伪造可解释性多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。