提升伪造视频检测公平性的同时,还能更好识别未知篡改手法。
Fair Deepfake Detectors Can Generalize

- 通过控制数据分布和模型能力等混淆因子,打通公平性与泛化性的因果路径。
- 在三个跨域基准上,公平性和泛化性能均优于现有顶尖检测器。
- 适合关注检测模型公平性与实际部署泛化能力的研究者与开发者。
深度伪造检测模型面临两大挑战:对未见过的篡改手法的泛化能力,以及不同人群间的公平性。然而,现有方法常显示这两项目标存在内在冲突。本文首次揭示并形式化定义了公平性与泛化性之间的因果关系。基于后门调整理论,我们证明控制混淆因子(数据分布与模型容量)可通过公平性干预实现更好的泛化。受此启发,提出可即插即用的DAID框架:(i) 基于逆倾向加权与子组特征归一化的性别/年龄感知数据重平衡,消除分布偏差;(ii) 利用新型对齐损失实现无敏感属性特征聚合,抑制身份相关信号。在三个跨域基准上,DAID在公平性与泛化性能上均显著优于多个先进检测器,验证了其理论基础与实际有效性。
原文摘要 · Abstract (English)
Deepfake detection models face two critical challenges: generalization to unseen manipulations and demographic fairness among population groups. However, existing approaches often demonstrate that these two objectives are inherently conflicting, revealing a trade-off between them. In this paper, we, for the first time, uncover and formally define a causal relationship between fairness and generalization. Building on the back-door adjustment, we show that controlling for confounders (data distribution and model capacity) enables improved generalization via fairness interventions. Motivated by this insight, we propose Demographic Attribute-insensitive Intervention Detection (DAID), a plug-and-play framework composed of: i) Demographic-aware data rebalancing, which employs inverse-propensity weighting and subgroup-wise feature normalization to neutralize distributional biases; and ii) Demographic-agnostic feature aggregation, which uses a novel alignment loss to suppress sensitive-attribute signals. Across three cross-domain benchmarks, DAID consistently achieves superior performance in both fairness and generalization compared to several state-of-the-art detectors, validating both its theoretical foundation and practical effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。