构建首个扩散模型生成的多模态数字人伪造数据集,推动深伪检测发展
Beyond Face Swapping: A Diffusion-Based Digital Human Benchmark for Multimodal Deepfake Detection
- 基于五种最新扩散模型与语音克隆技术构建大规模伪造视频数据集
- 数据集含6万视频(840万帧),真实感强,误认率高达68%
- 提出DigiShield检测模型,融合时空与跨模态特征,性能领先
近年来,基于扩散模型的数字人生成技术迅猛发展,对公共安全构成严峻威胁。与传统人脸操作不同,此类模型可通过多模态控制信号生成高度连贯逼真的视频,其灵活性与隐蔽性给现有检测方法带来巨大挑战。为此,我们构建了DigiFakeAV——首个基于扩散模型的大规模多模态数字人伪造数据集。该数据集涵盖五种前沿数字人生成方法及一种语音克隆技术,共包含60,000个视频(840万帧),覆盖多种族裔、肤色、性别与真实场景,显著提升数据多样性和真实性。用户研究显示,参与者对DigiFakeAV的误认率高达68%。同时,现有检测模型在该数据集上表现严重下降,凸显其挑战性。为应对该问题,我们提出DigiShield检测基线,通过联合建模视频的3D时空特征与音频的语义-声学特征,实现当前最优(SOTA)性能,并在其他数据集上展现强泛化能力。
原文摘要 · Abstract (English)
In recent years, the explosive advancement of deepfake technology has posed a critical and escalating threat to public security: diffusion-based digital human generation. Unlike traditional face manipulation methods, such models can generate highly realistic videos with consistency via multimodal control signals. Their flexibility and covertness pose severe challenges to existing detection strategies. To bridge this gap, we introduce DigiFakeAV, the new large-scale multimodal digital human forgery dataset based on diffusion models. Leveraging five of the latest digital human generation methods and a voice cloning method, we systematically construct a dataset comprising 60,000 videos (8.4 million frames), covering multiple nationalities, skin tones, genders, and real-world scenarios, significantly enhancing data diversity and realism. User studies demonstrate that the misrecognition rate by participants for DigiFakeAV reaches as high as 68%. Moreover, the substantial performance degradation of existing detection models on our dataset further highlights its challenges. To address this problem, we propose DigiShield, an effective detection baseline based on spatiotemporal and cross-modal fusion. By jointly modeling the 3D spatiotemporal features of videos and the semantic-acoustic features of audio, DigiShield achieves state-of-the-art (SOTA) performance on the DigiFakeAV and shows strong generalization on other datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。