arXiv:2508.21048cs.CVcs.AI2025-08被引 21

用模式感知推理提升深度伪造检测泛化能力

Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning

  • 引入模式感知推理,模拟人类取证思维过程
  • 在跨模型、新伪造技术等场景下准确率提升显著
  • 适合关注真实场景下伪造检测的从业者

深度伪造检测因现实场景中伪造内容的复杂性和动态性仍具挑战性。现有学术基准与工业实践存在严重脱节,通常训练数据同质化且测试图像质量低,制约了检测器的实际部署。为此,我们提出HydraFake数据集,模拟真实世界挑战,包含多样化伪造技术与真实环境伪造,采用严格的训练评估协议,覆盖未见模型架构、新兴伪造技术及新数据领域。基于此,我们提出Veritas,一种基于多模态大语言模型(MLLM)的深度伪造检测器。不同于传统链式思维(CoT),我们引入模式感知推理,包含‘规划’和‘自我反思’等关键推理模式,以模拟人类取证过程。进一步设计两阶段训练流程,将此类推理能力无缝内化至现有MLLM中。在HydraFake上的实验表明,尽管以往检测器在跨模型场景下表现良好,但在未见伪造技术和新数据域上表现不足。Veritas在多种分布外(OOD)场景下均取得显著提升,并能输出透明可信的检测结果。

原文摘要 · Abstract (English)

Deepfake detection remains a formidable challenge due to the complex and evolving nature of fake content in real-world scenarios. However, existing academic benchmarks suffer from severe discrepancies from industrial practice, typically featuring homogeneous training sources and low-quality testing images, which hinder the practical deployments of current detectors. To mitigate this gap, we introduce HydraFake, a dataset that simulates real-world challenges with hierarchical generalization testing. Specifically, HydraFake involves diversified deepfake techniques and in-the-wild forgeries, along with rigorous training and evaluation protocol, covering unseen model architectures, emerging forgery techniques and novel data domains. Building on this resource, we propose Veritas, a multi-modal large language model (MLLM) based deepfake detector. Different from vanilla chain-of-thought (CoT), we introduce pattern-aware reasoning that involves critical reasoning patterns such as "planning" and "self-reflection" to emulate human forensic process. We further propose a two-stage training pipeline to seamlessly internalize such deepfake reasoning capacities into current MLLMs. Experiments on HydraFake dataset reveal that although previous detectors show great generalization on cross-model scenarios, they fall short on unseen forgeries and data domains. Our Veritas achieves significant gains across different OOD scenarios, and is capable of delivering transparent and faithful detection outputs.

深度伪造多模态推理机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。