arXiv:2605.02567cs.CV2026-05被引 1

自动收集真实场景数据,让图像生成检测器持续进化

Automated In-the-Wild Data Collection for Continual AI Generated Image Detection

论文配图:Automated In-the-Wild Data Collection for Continual AI Generated Image Detection
图 1 · 摘自论文原文
  • 用事实核查文章自动构建真实世界数据集
  • 结合生成模型数据,检测准确率提升9.14%和8%
  • 适合需要长期更新检测能力的AI安全研究者

生成式人工智能的快速发展给可靠的图像生成检测带来挑战。现有检测器在分布偏移和面对新出现的生成模型时性能下降明显。本文提出一种以数据为中心的持续自适应框架,用于在动态环境中更新检测器。研究表明,真实世界数据与生成驱动数据均对检测器适应至关重要。我们设计了一种自动化、弱监督的数据构建流程,通过事实核查文章检索生成真实场景数据集。实验表明,即使少量生成驱动数据也能有效适应新兴模型;将其与真实世界数据结合,在持续学习框架下可实现鲁棒适应并缓解灾难性遗忘。在两个先进检测器上,平均准确率分别提升9.14%和8%。

原文摘要 · Abstract (English)

The rapid advancement of generative Artificial Intelligence (AI) has introduced significant challenges for reliable AI-generated image detection. Existing detectors often suffer from performance degradation under distribution shifts and when encountering newly emerging generative models. In this work, we propose a data-centric continual adaptation framework for updating detectors in evolving environments. We show that both in-the-wild data and generator-driven data are essential for adapting detectors. We introduce an automated, weakly supervised pipeline for constructing in-the-wild datasets through fact-check article retrieval. Additionally, we demonstrate that incorporating even a small amount of generator-driven data during training enables effective adaptation to newly emerging models, while combining it with in-the-wild data within a continual learning framework enables robust adaptation and mitigates catastrophic forgetting. Extensive experiments on two state-of-the-art detectors show significant improvements of +9.14% and +8% in average accuracy, respectively.

图像检测持续学习数据收集AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。