arXiv:2512.22303cs.CVcs.AI2025-12

提出可应对反伪造攻击的鲁棒检测模型,输出可信概率与可解释热图。

Attack-Aware Deepfake Detection under Counter-Forensic Manipulations

  • 双流结构:一通道抓语义,一通道提取伪造痕迹,轻量适配融合
  • 在多种攻击下保持高准确率,对抗重压缩和重着色时仍近满分排名
  • 生成面部区域聚焦的热图,无需精细标注,适合实际部署

本文提出一种攻击感知的深度伪造与图像取证检测器,旨在实现真实场景部署下的鲁棒性、校准概率与透明证据。方法采用红队训练结合测试时随机防御的两流架构:一流使用预训练主干提取语义内容,另一流提取取证残差,通过轻量级残差适配器融合分类;浅层特征金字塔风格头部在弱监督下生成篡改热图。红队训练每批次施加最坏情况的对抗性操作,包括JPEG重对齐与重压缩、重采样变形、去噪再赋纹、缝合平滑、微小色彩与伽马变化、社交应用转码等;测试时防御注入低开销抖动,如缩放、裁剪相位变化、轻微伽马变动及JPEG相位偏移,并聚合预测结果。热图通过人脸框掩码引导集中在面部区域,无需严格像素级标注。在标准深度伪造数据集与监控风格划分(低光、高压缩)上评估,报告了干净与受攻状态下的性能,包括AUC、最差情况准确率、可靠性、弃权质量与弱定位得分。结果表明,该方法在各类攻击下表现接近完美排序,校准误差极低,弃权风险小,且在重赋纹条件下退化可控,为攻击感知检测提供了模块化、数据高效、可实用的基准,具备校准概率与可行动热图。

原文摘要 · Abstract (English)

This work presents an attack-aware deepfake and image-forensics detector designed for robustness, well-calibrated probabilities, and transparent evidence under realistic deployment conditions. The method combines red-team training with randomized test-time defense in a two-stream architecture, where one stream encodes semantic content using a pretrained backbone and the other extracts forensic residuals, fused via a lightweight residual adapter for classification, while a shallow Feature Pyramid Network style head produces tamper heatmaps under weak supervision. Red-team training applies worst-of-K counter-forensics per batch, including JPEG realign and recompress, resampling warps, denoise-to-regrain operations, seam smoothing, small color and gamma shifts, and social-app transcodes, while test-time defense injects low-cost jitters such as resize and crop phase changes, mild gamma variation, and JPEG phase shifts with aggregated predictions. Heatmaps are guided to concentrate within face regions using face-box masks without strict pixel-level annotations. Evaluation on existing benchmarks, including standard deepfake datasets and a surveillance-style split with low light and heavy compression, reports clean and attacked performance, AUC, worst-case accuracy, reliability, abstention quality, and weak-localization scores. Results demonstrate near-perfect ranking across attacks, low calibration error, minimal abstention risk, and controlled degradation under regrain, establishing a modular, data-efficient, and practically deployable baseline for attack-aware detection with calibrated probabilities and actionable heatmaps.

深度伪造检测攻击感知可解释性弱监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。