构建视频伪造检测新基准,让模型学会识破动态造假痕迹。
Beyond Static Artifacts: A Forensic Benchmark for Video Deepfake Reasoning in Vision Language Models
- 将视频伪造分析转为多选题,分三层次评估模型能力。
- 在跨数据集测试中,微调后模型准确率提升显著。
- 适合研究视频真伪判断、多模态推理的学者使用。
当前视觉语言模型在识别静态伪造痕迹方面表现良好,但忽视了视频伪造中的时间不一致问题。为此,我们提出面向伪造推理的问答基准FAQ,将时序伪造分析转化为多选任务。FAQ采用三级层次体系:(1)面部感知,检验静态视觉异常识别能力;(2)时序伪造定位,要求模型跨帧定位动态伪造痕迹;(3)伪造推理,要求模型综合证据做出真伪判断。我们在FAQ上评估多种VLM,并构建指令微调数据集FAQ-IT。实验表明,经FAQ-IT微调的模型在同域与跨数据集检测任务中均达先进水平。消融实验证实,FAQ设计是推动模型时序推理能力的核心因素。
原文摘要 · Abstract (English)
Current Vision-Language Models (VLMs) for deepfake detection excel at identifying spatial artifacts but overlook a critical dimension: temporal inconsistencies in video forgeries. Adapting VLMs to reason about these dynamic cues remains a distinct challenge. To bridge this gap, we propose Forensic Answer-Questioning (FAQ), a large-scale benchmark that formulates temporal deepfake analysis as a multiple-choice task. FAQ introduces a three-level hierarchy to progressively evaluate and equip VLMs with forensic capabilities: (1) Facial Perception, testing the ability to identify static visual artifacts; (2) Temporal Deepfake Grounding, requiring the localization of dynamic forgery artifacts across frames; and (3) Forensic Reasoning, challenging models to synthesize evidence for final authenticity verdicts. We evaluate a range of VLMs on FAQ and generate a corresponding instruction-tuning set, FAQ-IT. Extensive experiments show that models fine-tuned on FAQ-IT achieve advanced performance on both in-domain and cross-dataset detection benchmarks. Ablation studies further validate the impact of our key design choices, confirming that FAQ is the driving force behind the temporal reasoning capabilities of these VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。