利用物理模型检测生成图像,无需训练即可准确识别。
Synthetic Image Detection via Spectral Gaps of QC-RBIM Nishimori Bethe-Hessian Operators
- 将图像特征建模为稀疏图,通过贝叶斯最优谱间隙判断真伪。
- 在真实图像上出现明显谱隙,生成图像则谱峰坍缩,准确率超94%。
- 无需标注数据或重训练,适用于新生成模型,适合媒体鉴真场景。
深度生成模型如GAN和扩散网络已能生成几乎无法与真实照片区分的图像,严重威胁媒体鉴伪与生物识别安全。监督检测方法对未见生成器或对抗后处理迅速失效,而依赖低级统计特征的无监督方法仍易出错。本文提出一种受物理启发、模型无关的检测方法,将合成图像识别视为稀疏加权图上的社区检测问题。使用预训练CNN提取图像特征并降维至32维,每个特征向量作为多边类型QC-LDPC图的节点。成对相似性经奈希莫里温度校准后转化为边耦合,构建随机键伊辛模型(RBIM),其贝叶斯-赫森谱在存在真实图像结构时呈现特征谱隙。合成图像破坏奈希莫里对称性,因而无此类谱隙。在Flickr-Faces-HQ和CelebA的真实图像与生成图像(来自GAN和扩散模型)的猫/狗、男/女二分类任务中验证:无需任何标注的合成数据或特征提取器重训练,检测准确率超过94%。谱分析显示真实图像集有多重分离谱隙,生成图像则谱峰坍缩。贡献包括:一种嵌入深度特征的新型LDPC图构造;奈希莫里温度下RBIM与贝叶斯-赫森谱的解析关联,提供贝叶斯最优检测准则;以及一种实用、无监督且对新生成架构鲁棒的合成图像检测器。未来工作将扩展至视频流与多类异常检测。
原文摘要 · Abstract (English)
The rapid advance of deep generative models such as GANs and diffusion networks now produces images that are virtually indistinguishable from genuine photographs, undermining media forensics and biometric security. Supervised detectors quickly lose effectiveness on unseen generators or after adversarial post-processing, while existing unsupervised methods that rely on low-level statistical cues remain fragile. We introduce a physics-inspired, model-agnostic detector that treats synthetic-image identification as a community-detection problem on a sparse weighted graph. Image features are first extracted with pretrained CNNs and reduced to 32 dimensions, each feature vector becomes a node of a Multi-Edge Type QC-LDPC graph. Pairwise similarities are transformed into edge couplings calibrated at the Nishimori temperature, producing a Random Bond Ising Model (RBIM) whose Bethe-Hessian spectrum exhibits a characteristic gap when genuine community structure (real images) is present. Synthetic images violate the Nishimori symmetry and therefore lack such gaps. We validate the approach on binary tasks cat versus dog and male versus female using real photos from Flickr-Faces-HQ and CelebA and synthetic counterparts generated by GANs and diffusion models. Without any labeled synthetic data or retraining of the feature extractor, the detector achieves over 94% accuracy. Spectral analysis shows multiple well separated gaps for real image sets and a collapsed spectrum for generated ones. Our contributions are threefold: a novel LDPC graph construction that embeds deep image features, an analytical link between Nishimori temperature RBIM and the Bethe-Hessian spectrum providing a Bayes optimal detection criterion; and a practical, unsupervised synthetic image detector robust to new generative architectures. Future work will extend the framework to video streams and multi-class anomaly detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。