真实平台环境让AI检测器失效,新框架揭示其脆弱性
The Deployment Gap in AI Media Detection: Platform-Aware and Visually Constrained Adversarial Evaluation
- 模拟平台压缩、截图等操作,用视觉合理扰动测试检测器
- 原AUC≈0.99的模型在真实场景下性能暴跌,误判率飙升
- 发现通用扰动存在,适合评估检测系统实际部署风险
近期AI媒体检测器在纯净实验室环境下表现接近完美,但其在真实部署条件下的鲁棒性尚未充分研究。实际上,生成图像在传播前会经历缩放、压缩、重编码和视觉修改等平台处理。我们指出这导致了实验室鲁棒性与现实可靠性之间的部署差距。本文提出一种平台感知的对抗评估框架,显式建模部署变换(如缩放、压缩、截图式失真),并将扰动限制在视觉合理的迷因风格条带内,而非全图噪声。在此威胁模型下,原本在干净设置中AUC≈0.99的检测器出现显著退化。单图平台感知攻击使AUC大幅下降,并实现高比例的假对真误判,即使有严格视觉约束。我们进一步证明,即便在局部条带约束下仍存在通用扰动,揭示输入间共享的脆弱方向。除准确率下降外,还观察到校准崩溃现象,检测器变得高度自信却错误。结果表明,干净条件下测得的鲁棒性严重高估了实际部署鲁棒性。我们主张将平台感知评估作为未来AI媒体安全基准的必要组成部分,并发布评估框架以推动标准化鲁棒性评测。
原文摘要 · Abstract (English)
Recent AI media detectors report near-perfect performance under clean laboratory evaluation, yet their robustness under realistic deployment conditions remains underexplored. In practice, AI-generated images are resized, compressed, re-encoded, and visually modified before being shared on online platforms. We argue that this creates a deployment gap between laboratory robustness and real-world reliability. In this work, we introduce a platform-aware adversarial evaluation framework for AI media detection that explicitly models deployment transforms (e.g., resizing, compression, screenshot-style distortions) and constrains perturbations to visually plausible meme-style bands rather than full-image noise. Under this threat model, detectors achieving AUC $\approx$ 0{.}99 in clean settings experience substantial degradation. Per-image platform-aware attacks reduce AUC to significantly lower levels and achieve high fake-to-real misclassification rates, despite strict visual constraints. We further demonstrate that universal perturbations exist even under localized band constraints, revealing shared vulnerability directions across inputs. Beyond accuracy degradation, we observe pronounced calibration collapse under attack, where detectors become confidently incorrect. Our findings highlight that robustness measured under clean conditions substantially overestimates deployment robustness. We advocate for platform-aware evaluation as a necessary component of future AI media security benchmarks and release our evaluation framework to facilitate standardized robustness assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。