用自我反思机制提升AI生成图像检测能力,持续进化更准更可靠。
Perception, Verdict, and Evolution: Hindsight-Driven Self-Refining Forensics Agent for AI-Generated Image Detection

- 分三域特征融合+大模型判断,实现逻辑化检测决策。
- 通过事后反思优化推理过程,自动生成高质量训练数据。
- 在16个生成器上均表现优异,适合需要高可信推理的场景。
生成模型的快速发展给现有深度伪造检测方法带来严峻挑战,尤其面对高度逼真的AI生成图像时。尽管多模态大语言模型(MLLM)在此任务中展现出潜力,但现有方法存在两大局限:对细微取证特征敏感度不足,且依赖前沿模型的静态合成监督,导致灵活性差、成本高。为此,我们提出ForeAgent,一种具有迭代自我演化的智能体式取证框架。首先,ForeAgent采用感知-判断架构,整合语义、空间与频域多视角特征,并以MLLM作为判断模块,融合信号生成逻辑化结论。其次,为实现持续自我改进,引入基于事后反思的自精炼策略,遵循采样-反思-演化范式:在训练样本上进行推理回放,以真实标签为事后认知,反思失败案例与低质量推理路径,生成更优推理轨迹;这些合成样本经双专家质量门控严格筛选后,用于持续微调。大量实验表明,ForeAgent在Chameleon基准上达到82.18%准确率(较AIDE提升16.41%),在涵盖16种生成器的AIGCDetect-Benchmark上平均准确率达93.3%。外部评估显示,其推理更具一致性与因果合理性,优于GPT-5和GPT-5-mini。
原文摘要 · Abstract (English)
The rapid advancement of generative models presents a significant challenge to existing deepfake detection methods, particularly given the widespread dissemination of highly realistic AI-generated images. Although Multimodal Large Language Models (MLLMs) show strong potential for this task, existing approaches suffer from two key limitations: insufficient sensitivity to fine-grained forensic artifacts and reliance on static synthetic supervision from frontier models, leading to limited flexibility and high-cost. To address these issues, we propose ForeAgent, an agentic forensics framework for AI-generated image detection with iterative self-evolution. First, ForeAgent adopts a Perception-Verdict architecture that aggregates multi-view cues spanning semantic, spatial, and frequency-domain features, and leverages an MLLM as a verdict module to fuse these signals for a logical-grounded verdict. Second, to enable continual self-improvement, we introduce a Hindsight-Driven Self-Refining strategy following a Sampling-Reflection-Evolution paradigm. The agent performs inference rollouts on training instances. Guided by ground-truth labels as hindsight, it reflects on failure cases and low-quality reasoning trajectories to regenerate higher-quality reasoning traces. These synthesized samples are then strictly filtered through a dual-expert quality gating module. ForeAgent continuously evolves via fine-tuning on self-curated high-quality samples. Extensive experiments demonstrate that ForeAgent achieves state-of-the-art performance on the Chameleon benchmark, reaching 82.18% accuracy (+16.41% over AIDE), and achieves 93.3% mean accuracy on AIGCDetect-Benchmark across 16 generators. In addition, external evaluation shows that ForeAgent produces more consistent and causally grounded reasoning compared to GPT-5 and GPT-5-mini.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。