提出新评估指标Cross-AUC,发现现有伪造检测模型跨数据集性能骤降。
Rethinking Cross-Domain Evaluation for Face Forgery Detection with Semantic Fine-grained Alignment and Mixture-of-Experts

- 用跨数据集对比真实与伪造样本,设计新评估指标Cross-AUC。
- 在Cross-AUC下,主流检测器性能下降超20%,暴露鲁棒性缺陷。
- 引入语义细粒度对齐与专家路由机制,提升面部区域伪造识别能力。
随着生成模型的快速发展,视觉伪造检测在社会与经济安全中愈发重要。现有面部伪造检测器泛化能力不足,关键原因在于缺乏合适评估指标:常用跨数据集AUC无法揭示检测分数在不同数据域间可能显著偏移的问题。为此,我们提出新指标Cross-AUC,通过对比一个数据集的真实样本与另一数据集的伪造样本(反之亦然)来计算跨数据集AUC。实验发现,基于Cross-AUC评估时,代表性检测器性能大幅下降,暴露出被忽视的鲁棒性问题。此外,我们提出SFAM框架,包含片级图像-文本对齐模块以增强CLIP对篡改痕迹的敏感性,以及面部区域专家路由模块,将不同面部区域特征分配给专用专家进行区域感知分析。在公开数据集上的大量定性与定量实验表明,该方法在多种指标下均优于当前最优方法。
原文摘要 · Abstract (English)
Nowadays, visual data forgery detection plays an increasingly important role in social and economic security with the rapid development of generative models. Existing face forgery detectors still can't achieve satisfactory performance because of poor generalization ability across datasets. The key factor that led to this phenomenon is the lack of suitable metrics: the commonly used cross-dataset AUC metric fails to reveal an important issue where detection scores may shift significantly across data domains. To explicitly evaluate cross-domain score comparability, we propose \textbf{Cross-AUC}, an evaluation metric that can compute AUC across dataset pairs by contrasting real samples from one dataset with fake samples from another (and vice versa). It is interesting to find that evaluating representative detectors under the Cross-AUC metric reveals substantial performance drops, exposing an overlooked robustness problem. Besides, we also propose the novel framework \textbf{S}emantic \textbf{F}ine-grained \textbf{A}lignment and \textbf{M}ixture-of-Experts (\textbf{SFAM}), consisting of a patch-level image-text alignment module that enhances CLIP's sensitivity to manipulation artifacts, and the facial region mixture-of-experts module, which routes features from different facial regions to specialized experts for region-aware forgery analysis. Extensive qualitative and quantitative experiments on the public datasets prove that the proposed method achieves superior performance compared with the state-of-the-art methods with various suitable metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。