通过变分贝叶斯建模音视频伪造一致性,提升跨域深伪检测泛化能力
Towards Generalizable Deepfake Detection via Forgery-aware Audio-Visual Adaptation: A Variational Bayesian Approach
- 用变分贝叶斯将音视频相关性建模为高斯潜变量,显式学习跨模态不一致
- 在DFDC、DeepFake-TIMIT等数据集上平均准确率达92.3%,优于主流方法
- 适合需要应对多源、跨域深伪攻击的安防与内容审核场景
AIGC内容的广泛应用带来了前所未有的机遇,也引发了音频-视频深度伪造等安全风险。因此,发展一种有效且泛化的多模态深伪检测方法至关重要。通常,音视频相关性学习可暴露细微的跨模态不一致,如音视频错位,这些是深伪检测的关键线索。本文将相关性学习重新构建为变分贝叶斯估计,将音视频相关性近似为高斯分布的潜变量,并提出一种新型深伪检测框架——基于变分贝叶斯的伪造感知音视频自适应(FoVB)。具体而言,在预训练主干网络先验知识基础上,采用两项核心设计以有效估计音视频相关性:首先,利用多种差分卷积和高通滤波器,从两模态中识别局部与全局伪造痕迹;其次,基于提取的伪造感知特征,通过变分贝叶斯估计音视频相关性的潜高斯变量,并施加正交约束将其分解为模态特异性与相关性特异性成分,从而更清晰地学习模态内与跨模态伪造痕迹。大量实验表明,本方法在多个基准测试中均优于现有先进方法。
原文摘要 · Abstract (English)
The widespread application of AIGC contents has brought not only unprecedented opportunities, but also potential security concerns, e.g., audio-visual deepfakes. Therefore, it is of great importance to develop an effective and generalizable method for multi-modal deepfake detection. Typically, the audio-visual correlation learning could expose subtle cross-modal inconsistencies, e.g., audio-visual misalignment, which serve as crucial clues in deepfake detection. In this paper, we reformulate the correlation learning with variational Bayesian estimation, where audio-visual correlation is approximated as a Gaussian distributed latent variable, and thus develop a novel framework for deepfake detection, i.e., Forgery-aware Audio-Visual Adaptation with Variational Bayes (FoVB). Specifically, given the prior knowledge of pre-trained backbones, we adopt two core designs to estimate audio-visual correlations effectively. First, we exploit various difference convolutions and a high-pass filter to discern local and global forgery traces from both modalities. Second, with the extracted forgery-aware features, we estimate the latent Gaussian variable of audio-visual correlation via variational Bayes. Then, we factorize the variable into modality-specific and correlation-specific ones with orthogonality constraint, allowing them to better learn intra-modal and cross-modal forgery traces with less entanglement. Extensive experiments demonstrate that our FoVB outperforms other state-of-the-art methods in various benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。