通过分析生成图像的离散码本分布差异,提升对自回归生成图像的检测精度。
$\bf{D^3}$QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection
- 基于码本频率分布偏差设计量化误差感知机制
- 在7种主流自回归模型上实现超90%检测准确率
- 适合图像伪造检测与AI生成内容监管场景
视觉自回归(AR)模型的兴起革新了图像生成技术,但也带来了合成图像检测的新挑战。与以往的GAN或扩散模型不同,AR模型通过离散标记预测生成图像,在向量量化表示上呈现出独特特征。本文提出一种离散分布差异感知量化误差(D³QE)方法,利用真实与伪造图像中码本存在的频率分布偏差进行检测。我们设计了一种动态码本频率统计融合的Transformer,将语义特征与量化误差潜在信息结合。为评估该方法,构建了覆盖7种主流视觉AR模型的综合性数据集ARForensics。实验表明,D³QE在不同AR模型间具有优异检测准确率和强泛化能力,且对真实世界干扰具有鲁棒性。
原文摘要 · Abstract (English)
The emergence of visual autoregressive (AR) models has revolutionized image generation while presenting new challenges for synthetic image detection. Unlike previous GAN or diffusion-based methods, AR models generate images through discrete token prediction, exhibiting both marked improvements in image synthesis quality and unique characteristics in their vector-quantized representations. In this paper, we propose to leverage Discrete Distribution Discrepancy-aware Quantization Error (D$^3$QE) for autoregressive-generated image detection that exploits the distinctive patterns and the frequency distribution bias of the codebook existing in real and fake images. We introduce a discrete distribution discrepancy-aware transformer that integrates dynamic codebook frequency statistics into its attention mechanism, fusing semantic features and quantization error latent. To evaluate our method, we construct a comprehensive dataset termed ARForensics covering 7 mainstream visual AR models. Experiments demonstrate superior detection accuracy and strong generalization of D$^3$QE across different AR models, with robustness to real-world perturbations. Code is available at \href{https://github.com/Zhangyr2022/D3QE}{https://github.com/Zhangyr2022/D3QE}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。