提出分层通道堆叠框架,让AI生成图像检测结果可解释。
Hierarchical Channel Stacking: A Structured Decision Framework for AI-Generated Image Detection

- 将CNN中间激活转为三级分层的60维结构表示
- 在多类生成器数据集上达86.7%准确率与宏F1
- 揭示不同生成类型在层级中的差异性贡献
许多合成图像检测器虽能准确判断,但缺乏决策过程的可解释性。本文提出分层通道堆叠(HCS)框架,将中间CNN激活转化为三个渐进深层阶段组织的60维结构化表示。HCS通过逐通道一级分类器和二级聚合器实现图像级预测,同时保留可分析的层级结构。在涵盖GAN与扩散模型生成器的基准测试中,HCS在预留测试集上取得86.7%准确率与86.7%宏F1。阶段消融实验表明,完整三阶段系统优于单阶段或两阶段变体,说明层级间包含互补预测信息。阶段贡献分析进一步显示,在所研究检测器设置下,伪造GAN图像与伪造扩散图像表现出显著不同的阶段贡献模式。这些结果表明,HCS不仅是紧凑检测器,更是一种用于研究合成图像检测如何跨表征层次整合证据的结构化框架。
原文摘要 · Abstract (English)
Many synthetic-image detectors produce accurate predictions but offer limited insight into how those decisions are formed. This paper introduces Hierarchical Channel Stacking (HCS), a compact framework for AI-generated image detection that converts intermediate CNN activations into a structured 60-dimensional representation organized across three progressively deeper backbone stages. HCS uses per-channel Level-1 classifiers and a Level-2 aggregator to produce image-level predictions while preserving explicit hierarchical structure for analysis. On a benchmark spanning GAN and diffusion generators, HCS achieves 86.7% accuracy and 86.7% macro-F1 on the held-out test set. Stage ablation shows that the full three-stage system outperforms reduced single-stage and two-stage variants, indicating that the hierarchy carries complementary predictive information. Stage-level contribution analysis further shows that, in the analyzed detector setting, fake GAN and fake diffusion images exhibit distinct stage-level contribution profiles. These results position HCS not simply as a compact detector, but as a structured framework for studying how synthetic-image detectors assemble evidence across representation levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。