通过修正真实数据分布偏差,提升深伪检测模型对未知伪造的泛化能力。
Towards Generalizable Deepfake Detection via Real Distribution Bias Correction
- 利用真实数据的群体分布和图像高斯性,构建双模块修正框架。
- 在跨域检测中达到当前最优效果,尤其在未见伪造类型上表现突出。
- 适合需要强泛化能力的深伪检测场景,如社交媒体内容安全审核。
为使深伪检测器泛化到未来未见的伪造类型,现有方法通常尝试用已有源域数据模拟不断演化的伪造形式。然而,仅凭有限先验样本预测无限可能的未来篡改行为是不可行的。为此,我们从两个互补视角挖掘真实数据的不变性:整个真实类别的固定群体分布,以及单个真实图像的固有高斯性。基于此,提出真实分布偏差修正(RDBC)框架,包含两个核心组件:真实群体分布估计模块与分布采样特征白化模块。前者利用真实样本的独立同分布(i.i.d.)特性,推导其统计量的正态分布形式,并仅用少量源域数据即可估计分布参数。后者基于真实数据的固有高斯性作为判别先验,通过采样白化操作放大真实与伪造样本间的高斯性差异。双模块协同作用使模型捕获真实世界中真实样本的本质属性,从而显著增强对未见目标域的泛化能力。大量实验表明,RDBC在域内与跨域深伪检测任务中均达到当前最优性能。
原文摘要 · Abstract (English)
To generalize deepfake detectors to future unseen forgeries, most existing methods attempt to simulate the dynamically evolving forgery types using available source domain data. However, predicting an unbounded set of future manipulations from limited prior examples is infeasible. To overcome this limitation, we propose to exploit the invariance of \textbf{real data} from two complementary perspectives: the fixed population distribution of the entire real class and the inherent Gaussianity of individual real images. Building on these properties, we introduce the Real Distribution Bias Correction (RDBC) framework, which consists of two key components: the Real Population Distribution Estimation module and the Distribution-Sampled Feature Whitening module. The former utilizes the independent and identically distributed (\iid) property of real samples to derive the normal distribution form of their statistics, from which the distribution parameters can be estimated using limited source domain data. Based on the learned population distribution, the latter utilizes the inherent Gaussianity of real data as a discriminative prior and performs a sampling-based whitening operation to amplify the Gaussianity gap between real and fake samples. Through synergistic coupling of the two modules, our model captures the real-world properties of real samples, thereby enhancing its generalizability to unseen target domains. Extensive experiments demonstrate that RDBC achieves state-of-the-art performance in both in-domain and cross-domain deepfake detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。