新模型UNITE可检测人脸、背景及全AI生成视频,突破传统检测局限。
Towards a Universal Synthetic Video Detector: From Face or Background Manipulations to Fully AI-Generated Content

- 用Transformer架构分析全帧特征,不依赖人脸存在
- 在跨数据集测试中对合成视频检测准确率超现有方法
- 适合需要泛化检测能力的安防与内容审核场景
现有DeepFake检测技术主要针对面部篡改,如换脸或唇同步。然而,文本到视频(T2V)和图像到视频(I2V)生成模型的发展使得完全由AI生成的内容和无缝背景修改成为可能,这挑战了以人脸为中心的检测方法,亟需更通用的解决方案。为此,我们提出通用篡改与合成视频识别网络(UNITE),其能捕捉全帧篡改,扩展至无脸、非人类主体及复杂背景修改场景。UNITE采用基于SigLIP-So400M基础模型提取的领域无关特征,结合变压器架构。由于涵盖面部/背景篡改及T2V/I2V内容的数据集有限,训练中引入任务无关数据。为缓解模型过度关注人脸的问题,加入注意力多样性(AD)损失,促进视频帧间空间注意力分布多样化。结合交叉熵损失后,检测性能在多种情境下显著提升。对比实验表明,UNITE在包含面部/背景篡改和全合成T2V/I2V视频的数据集上(跨数据集设置)优于现有最先进检测器,展现良好适应性与泛化能力。
原文摘要 · Abstract (English)
Existing DeepFake detection techniques primarily focus on facial manipulations, such as face-swapping or lip-syncing. However, advancements in text-to-video (T2V) and image-to-video (I2V) generative models now allow fully AI-generated synthetic content and seamless background alterations, challenging face-centric detection methods and demanding more versatile approaches. To address this, we introduce the \underline{U}niversal \underline{N}etwork for \underline{I}dentifying \underline{T}ampered and synth\underline{E}tic videos (\texttt{UNITE}) model, which, unlike traditional detectors, captures full-frame manipulations. \texttt{UNITE} extends detection capabilities to scenarios without faces, non-human subjects, and complex background modifications. It leverages a transformer-based architecture that processes domain-agnostic features extracted from videos via the SigLIP-So400M foundation model. Given limited datasets encompassing both facial/background alterations and T2V/I2V content, we integrate task-irrelevant data alongside standard DeepFake datasets in training. We further mitigate the model's tendency to over-focus on faces by incorporating an attention-diversity (AD) loss, which promotes diverse spatial attention across video frames. Combining AD loss with cross-entropy improves detection performance across varied contexts. Comparative evaluations demonstrate that \texttt{UNITE} outperforms state-of-the-art detectors on datasets (in cross-data settings) featuring face/background manipulations and fully synthetic T2V/I2V videos, showcasing its adaptability and generalizable detection capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。