用语义提示攻击破解AI生成内容检测器,让伪造图像逃过检测。
Fooling the Watchers: Breaking AIGC Detectors via Semantic Prompt Attacks
- 构建语法树与蒙特卡洛搜索结合的自动化提示生成框架。
- 在多款T2I模型上成功绕过开源和商业检测器,竞赛排名第一。
- 可生成高质量对抗数据集,助力检测系统训练与评估。
文本到图像(T2I)模型的发展使得生成逼真的人像成为可能,引发了身份滥用和AIGC检测器鲁棒性方面的严重担忧。本文提出一种自动化对抗提示生成框架,利用语法树结构与蒙特卡洛树搜索变体,系统探索语义提示空间。该方法生成多样且可控的提示,在多个T2I模型上均能持续规避开源及商用AIGC检测器。大量实验验证了其有效性,且在真实世界对抗AIGC检测竞赛中排名首位。除攻击场景外,该方法还可用于构建高质量对抗数据集,为训练和评估更鲁棒的AIGC检测与防御系统提供宝贵资源。
原文摘要 · Abstract (English)
The rise of text-to-image (T2I) models has enabled the synthesis of photorealistic human portraits, raising serious concerns about identity misuse and the robustness of AIGC detectors. In this work, we propose an automated adversarial prompt generation framework that leverages a grammar tree structure and a variant of the Monte Carlo tree search algorithm to systematically explore the semantic prompt space. Our method generates diverse, controllable prompts that consistently evade both open-source and commercial AIGC detectors. Extensive experiments across multiple T2I models validate its effectiveness, and the approach ranked first in a real-world adversarial AIGC detection competition. Beyond attack scenarios, our method can also be used to construct high-quality adversarial datasets, providing valuable resources for training and evaluating more robust AIGC detection and defense systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。