SPECTRA-Net通过多视角张量表示,实现AI生成图像的高精度跨域检测与可解释性。
SPECTRA-Net: Scalable Pipeline for Explainable Cross-domain Tensor Representations for AI-generated Images Detection

- 融合视觉基础模型、频谱分析与局部异常检测,构建多视角张量表征。
- 在WildFake、Chameleon等数据集上达到当前最优性能,跨域泛化能力强。
- 可定位伪造痕迹,适合需要可解释性的内容安全应用场景。
AI生成图像(AIGI)的快速扩散对数字信息真实性构成重大挑战。尽管人类观察者和现有检测模型难以跟上生成模型的迭代速度,但构建鲁棒、实时的检测系统已成为迫切需求。本文提出SPECTRA-Net,一个可扩展的可解释跨域张量表示管道,用于AIGI检测。该方法结合视觉基础模型(VFM)的全局语义特征、频谱分析、基于局部块的异常检测及统计描述符,融合互补的数据流。在包含WildFake、Chameleon、RRDataset在内的多个挑战性数据集上,SPECTRA-Net在域内与跨域设置下均表现出色,展现出高准确率与强泛化能力。该管道不仅提供高效的AIGI检测方案,还通过伪造区域定位实现可解释性,为真实世界中的可信内容验证开辟道路。
原文摘要 · Abstract (English)
The rapid proliferation of AI-generated images (AIGI) presents a significant challenge to digital information integrity. While human observers and existing detection models struggle to keep pace with the increasing sophistication of generative models, the need for robust, real-time detection systems has become critical. This paper introduces SPECTRA-Net, a scalable pipeline for explainable, cross-domain tensor representations for AIGI detection. Our approach leverages a multi-view representation of images, combining global semantic features from a Vision Foundation Model (VFM), spectral analysis, local patch-based anomaly detection, and statistical descriptors. By fusing these complementary data streams, SPECTRA-Net achieves state-of-the-art performance in both in-domain and cross-domain settings, demonstrating high accuracy and generalization capabilities across a wide range of challenging datasets, including WildFake, Chameleon, and RRDataset. The proposed pipeline not only provides a robust solution for AIGI detection but also offers explainability through artifact localization, paving the way for more trustworthy and reliable content verification in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。