用视觉语言模型提升深度伪造检测的准确性和可解释性
MARE: Multimodal Alignment and Reinforcement for Explainable Deepfake Detection via Vision-Language Models
- 结合人类反馈强化学习,让模型生成与图像对齐的推理文本
- 通过伪造解耦模块捕捉面部语义中的伪造痕迹,提升检测能力
- 适合关注可解释性检测和生成式内容安全的研究者
深度伪造检测是应对恶意内容传播的关键课题,现有方法多将其视为分类或空间定位问题。生成模型的快速演进对检测技术提出新要求。本文提出基于视觉语言模型的多模态对齐与强化方法MARE,旨在提升视觉语言模型在深度伪造检测与推理中的准确性与可靠性。MARE设计了综合奖励函数,引入人类反馈强化学习(RLHF),激励模型生成符合人类偏好的文本-空间对齐推理内容。此外,MARE引入伪造解耦模块,从高层面部语义中捕捉内在伪造痕迹,从而增强真实性检测能力。我们在生成的推理内容上进行了全面评估,定量与定性结果均表明,MARE在准确性和可靠性方面达到当前最优水平。
原文摘要 · Abstract (English)
Deepfake detection is a widely researched topic that is crucial for combating the spread of malicious content, with existing methods mainly modeling the problem as classification or spatial localization. The rapid advancements in generative models impose new demands on Deepfake detection. In this paper, we propose multimodal alignment and reinforcement for explainable Deepfake detection via vision-language models, termed MARE, which aims to enhance the accuracy and reliability of Vision-Language Models (VLMs) in Deepfake detection and reasoning. Specifically, MARE designs comprehensive reward functions, incorporating reinforcement learning from human feedback (RLHF), to incentivize the generation of text-spatially aligned reasoning content that adheres to human preferences. Besides, MARE introduces a forgery disentanglement module to capture intrinsic forgery traces from high-level facial semantics, thereby improving its authenticity detection capability. We conduct thorough evaluations on the reasoning content generated by MARE. Both quantitative and qualitative experimental results demonstrate that MARE achieves state-of-the-art performance in terms of accuracy and reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。