用合成数据提升深度伪造检测的公平性泛化能力
Data-Driven Fairness Generalization for Deepfake Detection
- 通过生成跨种族/性别的合成数据增强训练集多样性
- 在跨数据集评估中显著优于现有方法,公平性指标更稳定
- 适合关注模型公平性与实际部署可靠性的研究者
尽管深度伪造检测研究取得进展,但训练数据中的偏差导致检测性能在不同人种和性别群体间存在差异,可能引发不公平识别。传统基于公平损失函数的方法在未见数据集上表现不佳,公平性泛化仍是挑战。本文提出一种数据驱动框架,利用合成数据与优化策略提升公平性泛化能力。通过构建涵盖多种族、性别特征的合成样本,实现平衡且具代表性的训练数据。结合损失曲率感知优化与多任务学习,有效提升模型在同数据集内与跨数据集评估中的公平性。在多个基准数据集上的实验证明,该方法在跨数据集评估中显著优于现有技术,验证了合成数据在实现公平性泛化中的潜力,为深度伪造检测提供了更可靠的解决方案。
原文摘要 · Abstract (English)
Despite the progress made in deepfake detection research, recent studies have shown that biases in the training data for these detectors can result in varying levels of performance across different demographic groups, such as race and gender. These disparities can lead to certain groups being unfairly targeted or excluded. Traditional methods often rely on fair loss functions to address these issues, but they under-perform when applied to unseen datasets, hence, fairness generalization remains a challenge. In this work, we propose a data-driven framework for tackling the fairness generalization problem in deepfake detection by leveraging synthetic datasets and model optimization. Our approach focuses on generating and utilizing synthetic data to enhance fairness across diverse demographic groups. By creating a diverse set of synthetic samples that represent various demographic groups, we ensure that our model is trained on a balanced and representative dataset. This approach allows us to generalize fairness more effectively across different domains. We employ a comprehensive strategy that leverages synthetic data, a loss sharpness-aware optimization pipeline, and a multi-task learning framework to create a more equitable training environment, which helps maintain fairness across both intra-dataset and cross-dataset evaluations. Extensive experiments on benchmark deepfake detection datasets demonstrate the efficacy of our approach, surpassing state-of-the-art approaches in preserving fairness during cross-dataset evaluation. Our results highlight the potential of synthetic datasets in achieving fairness generalization, providing a robust solution for the challenges faced in deepfake detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。