arXiv:2603.21526cs.CV2026-03被引 1

VIGIL通过分步检测人脸部位,提升伪造视频识别的准确性和可解释性。

VIGIL: Part-Grounded Structured Reasoning for Generalizable Deepfake Detection

  • 先规划需检查的面部区域,再独立验证各区域证据
  • 在跨数据集测试中超越现有方法,在5级泛化层级上均表现优异
  • 适合需要高可信度与可解释性的深度伪造检测场景

多模态大语言模型(MLLM)为可解释的深度伪造检测提供了前景,但当前方法将证据生成与篡改定位合并处理,模糊了真实观察与幻觉解释的界限,导致结论不可靠。为此,我们提出VIGIL,一种受专家法证实践启发的部件中心结构化框架,采用‘规划-检验’流程:模型先基于全局视觉线索规划应检查的面部部件,再对每个部件使用独立来源的取证证据进行分析。阶段门控注入机制仅在检验阶段引入部件级取证证据,确保部件选择由模型自身感知驱动,避免外部信号干扰。我们进一步设计渐进式三阶段训练范式,其中强化学习阶段采用部件感知奖励,以强制满足解剖学合理性与证据-结论一致性。为实现严格的泛化能力评估,我们构建了OmniFake,一个分层5级基准,模型仅在3个基础生成器上训练,逐步测试至真实社交媒体数据。大量实验表明,VIGIL在OmniFake及跨数据集评估中,始终优于专家检测器和现有MLLM方法,覆盖所有泛化层级。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) offer a promising path toward interpretable deepfake detection by generating textual explanations. However, the reasoning process of current MLLM-based methods combines evidence generation and manipulation localization into a unified step. This combination blurs the boundary between faithful observations and hallucinated explanations, leading to unreliable conclusions. Building on this, we present VIGIL, a part-centric structured forensic framework inspired by expert forensic practice through a plan-then-examine pipeline: the model first plans which facial parts warrant inspection based on global visual cues, then examines each part with independently sourced forensic evidence. A stage-gated injection mechanism delivers part-level forensic evidence only during examination, ensuring that part selection remains driven by the model's own perception rather than biased by external signals. We further propose a progressive three-stage training paradigm whose reinforcement learning stage employs part-aware rewards to enforce anatomical validity and evidence--conclusion coherence. To enable rigorous generalizability evaluation, we construct OmniFake, a hierarchical 5-Level benchmark where the model, trained on only three foundational generators, is progressively tested up to in-the-wild social-media data. Extensive experiments on OmniFake and cross-dataset evaluations demonstrate that VIGIL consistently outperforms both expert detectors and concurrent MLLM-based methods across all generalizability levels.

深度伪造检测可解释性多模态模型结构化推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。