arXiv:2606.07034cs.CV2026-06中稿 · ICML被引 1

提取AI图像伪造特征并跨模型迁移,提升检测泛化能力

ForensicConcept: Transferable Forensic Concepts for AIGI Detection

论文配图:ForensicConcept: Transferable Forensic Concepts for AIGI Detection
图 1 · 摘自论文原文
  • 通过Transformer归因定位关键判别区域,聚类生成可解释的伪造概念码本
  • 在多个数据集上检测准确率显著优于现有方法,尤其对未知生成器有效
  • 首次实现扩散模型特征向目标模型迁移,适合需要可解释检测的场景

当前AI生成图像检测器在分布内数据上表现优异,但在未见过的生成器上性能急剧下降。主要瓶颈在于检测器的黑箱特性——无法揭示其决策依据。我们提出ForensicConcept框架,从检测器中提取显式的伪造概念,并实现跨主干网络的迁移。该方法利用Transformer归因定位决策关键区域,将其聚类形成紧凑的概念码本,并通过概念对齐投影生成可审计的证据输出。受先前研究启发,DINO表示能引导扩散生成且与扩散特征存在概念级对应,我们引入基于CleanDIFT扩散特征的生成溯源参考,并通过邻域结构一致性(CKNNA)量化主干网络与生成特征的对齐程度。进一步提出概念码本注入机制,将扩散模型导出的概念迁移至目标主干网络。在GenImage、GAN家族和Chameleon基准上的实验显示,相比已有方法持续提升。同时发现CKNNA对齐度可预测迁移效果,为不同主干网络产生可迁移伪造证据的能力差异提供了理论解释。

原文摘要 · Abstract (English)

AI-generated image detectors achieve high accuracy on in-distribution data but often fail on unseen generators. A key obstacle to understanding this failure is the black-box nature of current detectors: they do not reveal which evidence drives their decisions. We propose ForensicConcept, a framework that extracts explicit forensic concepts from detectors and enables their transfer across backbones. Our method localizes decision-critical patches via Transformer attribution, clusters them into a compact concept codebook, and uses a concept-aligned projection to produce auditable evidence readouts. Motivated by prior studies showing that DINO representations can guide diffusion generation and exhibit concept-level correspondence with diffusion features, we introduce a generation-trace reference based on CleanDIFT diffusion features and quantify backbone-trace alignment via neighborhood-structure consistency (CKNNA). We further propose concept codebook injection to transfer diffusion-derived concepts into target backbones. Experiments on GenImage, GAN-family, and Chameleon benchmarks show consistent improvements over prior methods. We also find that CKNNA alignment predicts transfer effectiveness, providing a principled explanation for why some backbones yield more transferable forensic evidence than others.

AIGI检测可解释性特征迁移扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。