现有AI绘画检测器在新生成模型下失效,需构建更鲁棒的防御体系。
Robustness of AI-Art Detectors under Generator Shift

- 用反向提示生成十种风格图像,测试检测器跨架构泛化能力
- 基于U-Net的检测器在新模型上误判率飙升,人类作品误报仍低
- CLIP ViT-L/14表现最优,但误检图像激活区域模糊
文本到图像生成模型快速发展,现代扩散变换器架构生成的图像越来越难以与人类创作区分。这引发了版权保护、虚假信息、欺诈、冒名顶替和数字内容真实性等重大问题。现有AI绘画检测器大多在同一代型模型上训练和评估,对新型架构的鲁棒性研究不足。本研究基于稳定扩散3.5中型(SD3.5m)艺术数据集,通过反向提示未见的人类艺术品样本,分析生成器迁移影响。五个检测器在基于U-Net的潜在扩散艺术图像上训练,并在零样本跨生成器设置下于SD3.5m数据集上评估。深度学习模型在分布内表现良好,但在生成器迁移下性能下降,大量SD3.5m图像被错误分类为人类作品,而人类假阳性保持低位。CLIP ViT-L/14整体表现最佳,但梯度激活图显示误检样本激活区域较弱且分散。这些发现揭示了当前检测器的泛化差距,强调应将检测器作为多层次防御体系的一部分,以应对快速演进的生成架构。
原文摘要 · Abstract (English)
Text-to-image generative models have advanced rapidly, with modern Diffusion Transformer architectures producing images that are increasingly difficult to distinguish from human-created artwork. This development has raised significant concerns regarding copyright protection, misinformation, fraud, impersonation, and the authenticity of digital content. Most AI-art detectors are trained and evaluated on the same generator family, leaving robustness to newer architectures underexplored. In this chapter, we analyze generator shift based on a Stable Diffusion 3.5 Medium (SD3.5m) artwork dataset spanning ten art styles through reverse prompting of held-out human artwork samples. Five detectors are trained on U-Net-based latent diffusion artwork and evaluated in a zero-shot cross-generator setting on the SD3.5m dataset. Deep learning models perform strongly in-distribution but degrade under generator shift, misclassifying many SD3.5m images as human while human false positives remain low. The CLIP ViT-L/14 model performs best overall, while Grad-CAM analysis reveals weaker and more diffuse activation on false negatives. These findings highlight a generalization gap in current AI-art detectors and motivate the development of detectors as one component of a layered defense that remains reliable across rapidly evolving generative architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。