arXiv:2606.19184cs.CVcs.LG2026-06

提出新评估指标,更真实反映深伪检测器在跨域场景下的性能

When AUC Misleads: Polarization-Aware Evaluation of Deepfake Detectors under Domain Shift

论文配图:When AUC Misleads: Polarization-Aware Evaluation of Deepfake Detectors under Domain Shift
图 1 · 摘自论文原文
  • 用跨数据集AUC与预测极化度结合评估检测器泛化能力
  • 实验显示传统AUC在跨域时可能高估性能,新指标更可靠
  • 适合关注检测器实际应用效果的研究者和开发者

生成式AI(如扩散模型和换脸工具)的进步使得深伪内容日益逼真,带来金融诈骗和非自愿露骨内容等现实危害。为此,深伪检测成为研究热点,近期方法更注重对未见篡改手法的泛化能力。当前通常在多个数据集上分别计算受试者工作特征曲线下面积(AUC)进行评估,但这种做法无法反映真实场景中多种数据源混合、伪造类型多样的情况。为此,本文提出一种新指标——跨数据集AUC(Cross-AUC),通过加权平均各领域AUC并引入预测得分分布极化程度(以Wasserstein距离衡量)来评估模型在域偏移下的鲁棒性。该指标不仅更真实地反映检测器在跨域条件下的表现,还具有可解释性,能揭示性能下降的原因。在七个基准数据集上的实验证明其实际意义。

原文摘要 · Abstract (English)

Recent advances in generative AI, such as diffusion models and face-swapping tools, have enabled the creation of highly realistic deepfakes, leading to real-world harms including financial fraud and non-consensual explicit content. In response, deepfake detection has become an active research area, with recent methods increasingly focusing on improving generalization to unseen manipulations. This is typically evaluated using the Area Under the ROC Curve (AUC) measured separately across multiple datasets. However, such an evaluation fails to reflect real-world scenarios where detectors face a mixture of data sources and varying artifact types. To address this limitation, we introduce a novel metric, Cross-dataset AUC (Cross-AUC) that averages per-domain AUCs with a measure of prediction polarization for taking into account the robustness to domain shift. The polarization extent is quantified by the Wasserstein Distance between class score distributions. Cross-AUC not only assesses the generalization capabilities of deepfake detectors under domain shifts more realistically, but it is also interpretable as it better explains the reason behind a drop in performance. Experiments performed on seven benchmark datasets demonstrate its practical relevance.

深伪检测评估指标域偏移AUC

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。