新方法让检测AI生成图像更适应新场景和新技术。
SAIDO: Generalizable Detection of AI-Generated Images via Scene-Aware and Importance-Guided Dynamic Optimization in Continual Learning
- 用视觉大模型动态识别场景,为每类场景分配专用检测模块。
- 在持续学习中减少遗忘,检测错误率下降44.22%,遗忘率降40.57%。
- 适合需要长期更新、应对新生成技术的图像安全检测系统。
图像生成技术的滥用引发安全担忧,推动了AI生成图像检测方法的发展。然而,泛化能力已成为关键挑战:现有方法难以适应现实场景中不断出现的新生成技术与内容类型。为此,我们提出一种基于持续学习的场景感知与重要性引导动态优化检测框架(SAIDO)。具体地,设计了基于场景感知的专家模块(SAEM),利用视觉大模型动态识别并引入新场景,为每个场景分配独立的专家模块,从而更好捕捉场景特异性伪造特征,提升跨场景泛化能力。为缓解多生成方法学习过程中的灾难性遗忘问题,提出重要性引导的动态优化机制(IDOM),通过重要性引导的梯度投影策略优化每个神经元,实现模型可塑性与稳定性的有效平衡。大量持续学习实验表明,该方法在稳定性与可塑性上均优于当前最先进方法,平均检测误差率降低44.22%,遗忘率降低40.57%。在开放世界数据集上,平均检测准确率较当前最先进方法提升9.47%。
原文摘要 · Abstract (English)
The widespread misuse of image generation technologies has raised security concerns, driving the development of AI-generated image detection methods. However, generalization has become a key challenge and open problem: existing approaches struggle to adapt to emerging generative methods and content types in real-world scenarios. To address this issue, we propose a Scene-Aware and Importance-Guided Dynamic Optimization detection framework with continual learning (SAIDO). Specifically, we design Scene-Awareness-Based Expert Module (SAEM) that dynamically identifies and incorporates new scenes using VLLMs. For each scene, independent expert modules are dynamically allocated, enabling the framework to capture scene-specific forgery features better and enhance cross-scene generalization. To mitigate catastrophic forgetting when learning from multiple image generative methods, we introduce Importance-Guided Dynamic Optimization Mechanism (IDOM), which optimizes each neuron through an importance-guided gradient projection strategy, thereby achieving an effective balance between model plasticity and stability. Extensive experiments on continual learning tasks demonstrate that our method outperforms the current SOTA method in both stability and plasticity, achieving 44.22\% and 40.57\% relative reductions in average detection error rate and forgetting rate, respectively. On open-world datasets, it improves the average detection accuracy by 9.47\% compared to the current SOTA method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。