用多智能体协作分析图像真伪,准确率达97.05%
From Evidence to Verdict: An Agent-Based Forensic Framework for AI-Generated Image Detection
- 通过多智能体协作进行多阶段取证分析
- 在6000张图像上达到97.05%检测准确率
- 适合需要可解释性与跨模型泛化的检测场景
AI生成图像的快速演进对信息完整性和媒体真实性构成日益严峻挑战。现有检测方法在鲁棒性、可解释性及跨多种生成模型的泛化能力方面存在局限,尤其依赖单一视觉证据时。我们提出AIFo(基于智能体的图像取证)框架,一种无需训练的检测方法,将AI生成图像识别建模为多阶段取证分析过程,通过多智能体协作实现。该框架整合反向图像搜索、元数据提取、预训练分类器和视觉语言模型分析等工具,并通过结构化多智能体辩论机制解决证据不足或冲突问题。可选的记忆增强模块支持引入历史案例信息。我们在涵盖6000张图像的基准上评估,覆盖受控实验室环境与复杂真实场景,AIFo取得97.05%准确率,持续优于传统分类器与强视觉语言模型基线。结果证明,基于智能体的过程推理在AI生成图像检测中具有显著有效性。
原文摘要 · Abstract (English)
The rapid evolution of AI-generated images poses growing challenges to information integrity and media authenticity. Existing detection approaches face limitations in robustness, interpretability, and generalization across diverse generative models, particularly when relying on a single source of visual evidence. We introduce AIFo (Agent-based Image Forensics), a training-free framework that formulates AI-generated image detection as a multi-stage forensic analysis process through multi-agent collaboration. The framework integrates a set of forensic tools, including reverse image search, metadata extraction, pre-trained classifiers, and vision-language model analysis, and resolves insufficient or conflicting evidence through a structured multi-agent debate mechanism. An optional memory-augmented module further enables the framework to incorporate information from historical cases. We evaluate AIFo on a benchmark of 6,000 images spanning controlled laboratory settings and challenging real-world scenarios, where it achieves 97.05% accuracy and consistently outperforms traditional classifiers and strong vision-language model baselines. These findings demonstrate the effectiveness of agent-based procedural reasoning for AI-generated image detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。