arXiv:2506.11031cs.LGcs.AI2025-06

用提示引导大模型检测AI生成图像,零样本效果提升24%。

Prefill-Guided Thinking for zero-shot detection of AI-generated images

  • 用特定提示词引导视觉语言模型推理,提升检测能力。
  • 在16种生成器图像上,宏平均F1最高提升24%。
  • 适合关注AI内容安全与零样本检测的研究者。

传统监督方法依赖大量标注数据,难以泛化到新生成器。我们探索预训练视觉语言模型(VLM)在零样本场景下检测AI生成图像的可行性。在涵盖16种先进图像生成器的三类基准测试中(人脸、物体、动物),现成VLM表现不佳。我们提出预填充引导思维(Prefill-Guided Thinking, PGT):通过在响应前添加提示如“让我们分析风格和合成伪影”,显著提升三款主流开源VLM的性能,宏平均F1最高提升24%。通过追踪生成过程中的置信度,发现预填充可缓解部分模型的早期过度自信(类似达克效应),从而改善检测效果。

原文摘要 · Abstract (English)

Traditional supervised methods for detecting AI-generated images depend on large, curated datasets for training and fail to generalize to novel, out-of-domain image generators. As an alternative, we explore pre-trained Vision-Language Models (VLMs) for zero-shot detection of AI-generated images. We evaluate VLM performance on three diverse benchmarks encompassing synthetic images of human faces, objects, and animals produced by 16 different state-of-the-art image generators. While off-the-shelf VLMs perform poorly on these datasets, we find that prefilling responses effectively guides their reasoning -- a method we call Prefill-Guided Thinking (PGT). In particular, prefilling a VLM response with the phrase "Let's examine the style and the synthesis artifacts" improves the Macro F1 scores of three widely used open-source VLMs by up to 24%. We analyze this improvement in detection by tracking answer confidence during response generation. For some models, prefills counteract early overconfidence -- akin to mitigating the Dunning-Kruger effect -- leading to better detection performance.

AI检测视觉语言模型零样本提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。