arXiv:2410.23623cs.CV2024-10NeurIPS被引 61

用多模态特征检测扩散生成视频,提升对未知伪造内容的识别能力。

On Learning Multi-Modal Forgery Representation for Diffusion Generated Video Detection

  • 基于大模型多模态空间生成伪造特征表示,增强泛化能力。
  • 在DVF数据集上准确率达98.7%,优于现有方法。
  • 适合视频安全、AI内容监管领域研究人员使用。

大量由扩散模型生成的视频对信息安全性与真实性构成威胁,亟需高效的内容检测手段。现有视频级检测算法多聚焦于人脸伪造,难以识别语义多样化的扩散生成内容。为此,本文提出多模态检测算法MM-Det,利用大模型(LMMs)在多模态空间中生成多模态伪造表示(MMFR),显著提升对未见伪造内容的检测能力。同时,引入帧内与跨帧注意力机制(IAFA)进行时空特征增强,并设计动态融合策略优化伪造表征。此外,构建了覆盖多种伪造类型的扩散视频数据集DVF。实验表明,MM-Det在DVF上达到98.7%的准确率,性能领先当前方法。源代码与数据集已开源。

原文摘要 · Abstract (English)

Large numbers of synthesized videos from diffusion models pose threats to information security and authenticity, leading to an increasing demand for generated content detection. However, existing video-level detection algorithms primarily focus on detecting facial forgeries and often fail to identify diffusion-generated content with a diverse range of semantics. To advance the field of video forensics, we propose an innovative algorithm named Multi-Modal Detection(MM-Det) for detecting diffusion-generated videos. MM-Det utilizes the profound perceptual and comprehensive abilities of Large Multi-modal Models (LMMs) by generating a Multi-Modal Forgery Representation (MMFR) from LMM's multi-modal space, enhancing its ability to detect unseen forgery content. Besides, MM-Det leverages an In-and-Across Frame Attention (IAFA) mechanism for feature augmentation in the spatio-temporal domain. A dynamic fusion strategy helps refine forgery representations for the fusion. Moreover, we construct a comprehensive diffusion video dataset, called Diffusion Video Forensics (DVF), across a wide range of forgery videos. MM-Det achieves state-of-the-art performance in DVF, demonstrating the effectiveness of our algorithm. Both source code and DVF are available at https://github.com/SparkleXFantasy/MM-Det.

视频伪造检测扩散模型多模态学习AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。