让AI像专家一样分步推理,识别假图更准更可信。
ClueAegis: Heuristic-to-Reasoning Cognitive-skill Learning for Unified Evidence-based Synthetic Image Detection

- 从感知线索出发,选最优检测技能再推理
- 在多个数据集上超越现有方法,泛化性更强
- 输出可解释的推理过程,适合安全审查场景
生成模型的快速发展使合成图像日益逼真,给可靠检测带来挑战。现有方法多为端到端分类或单一推理模式,难以建模结构化的司法取证推理和异构视觉证据。本文从认知视角重新审视合成图像检测,提出一种‘启发式到推理’的认知技能学习框架,用于基于证据的司法分析。输入图像后,框架先提取感知线索,选择最优检测技能,再进行技能条件下的证据提取与决策推理。为此构建了ClueAegis-Bench基准,将检测任务分解为显式标注的司法认知技能,实现超越二分类的结构化评估。基于此,提出ClueAegis(Cognitive-skill Learning for Unified Evidence-based Synthetic Image Detection),一个两阶段代理式框架,先进行启发式技能选择,再通过技能条件工具链执行证据引导推理。该设计将合成图像检测重构为可配置的多技能推理流程,贯通感知、技能选择与司法推理。大量实验表明,ClueAegis达到当前最佳性能,同时提升跨域泛化与鲁棒性,并提供透明的推理轨迹与结构化证据,为传统端到端检测器提供更具可解释性的替代方案。
原文摘要 · Abstract (English)
The rapid advancement of generative models has made synthetic images increasingly realistic, challenging reliable detection. Existing methods are often limited to end-to-end classification or monolithic reasoning, and thus fail to model structured forensic reasoning and heterogeneous visual evidence. We revisit synthetic image detection from a cognitive perspective and propose a \textit{Heuristic-to-Reasoning} cognitive skill learning framework for evidence-based forensic analysis. Given an input image, our framework first extracts heuristic perceptual clues, selects the optimal forensic skill, and then performs skill-conditioned reasoning for evidence extraction and decision making. To support this paradigm, we introduce \textbf{ClueAegis-Bench}, which decomposes synthetic image detection into explicitly annotated forensic cognitive skills for structured evaluation beyond binary classification. Based on this benchmark, we propose \textbf{ClueAegis} (\underline{C}ognitive-skill \underline{L}earning for \underline{U}nified \underline{E}vidence-based Synthetic Image Detection), a two-stage agentic framework that conducts heuristic skill selection followed by evidence-guided reasoning through skill-conditioned toolchains. This design reformulates synthetic image detection as a configurable multi-skill reasoning process that bridges perception, skill selection, and forensic reasoning. Extensive experiments show that ClueAegis achieves state-of-the-art performance while improving cross-domain generalization and robustness. It also provides transparent reasoning trajectories and structured forensic evidence, offering a more explainable alternative to conventional end-to-end detectors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。