arXiv:2503.14421cs.CVcs.AI2025-03中稿 · WACV 2026被引 18

首个可解释的视频深度伪造检测数据集,助力精准识别伪造痕迹。

ExDDV: A New Dataset for Explainable Deepfake Detection in Video

  • 构建5.4K视频数据集,含人工标注的文字说明与点击定位。
  • 实验表明文本与点击监督缺一不可,才能实现精准定位与解释。
  • 适合研究可解释性深度伪造检测、AI安全与可信AI的学者使用。

随着生成视频的真实度和质量不断提升,人类难以辨别深度伪造内容,愈发依赖自动检测工具。然而,现有检测器易出错且决策过程不透明,使人们易受深度伪造欺诈和虚假信息侵害。为此,我们提出ExDDV,首个面向视频可解释深度伪造检测的数据集与基准。ExDDV包含约5.4K条真实与伪造视频,均经人工标注文字描述(解释伪造痕迹)和点击位置(定位痕迹)。我们在ExDDV上评估多种视觉-语言模型,采用不同微调与上下文学习策略。结果表明,仅靠文本或点击监督均不足,二者结合才能训练出鲁棒的可解释模型,有效定位并描述伪造特征。相关数据集与代码已开源:https://github.com/vladhondru25/ExDDV。

原文摘要 · Abstract (English)

The ever growing realism and quality of generated videos makes it increasingly harder for humans to spot deepfake content, who need to rely more and more on automatic deepfake detectors. However, deepfake detectors are also prone to errors, and their decisions are not explainable, leaving humans vulnerable to deepfake-based fraud and misinformation. To this end, we introduce ExDDV, the first dataset and benchmark for Explainable Deepfake Detection in Video. ExDDV comprises around 5.4K real and deepfake videos that are manually annotated with text descriptions (to explain the artifacts) and clicks (to point out the artifacts). We evaluate a number of vision-language models on ExDDV, performing experiments with various fine-tuning and in-context learning strategies. Our results show that text and click supervision are both required to develop robust explainable models for deepfake videos, which are able to localize and describe the observed artifacts. Our novel dataset and code to reproduce the results are available at https://github.com/vladhondru25/ExDDV.

深度伪造可解释性视频检测数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。