通过分析生成图像的低秩残差特征,精准识别AI伪造图像。
LoRC: Detecting AI-Generated Images via Low-Rank Collapse in Semantic Residuals

- 从语义残差的几何结构出发,发现生成模型共有的低秩坍缩现象。
- 在39个未见生成器上达到97.0%准确率,平均提升7.0%。
- 不依赖具体模型架构,适合真实场景下大规模检测。
现代生成模型能忠实建模宏观语义,生成高度逼真的合成图像,因此决定性取证线索存在于细微的非语义视觉差异中。本文从几何角度重新审视AI生成图像(AIGI)检测,发现一种与架构无关的特征:现代生成器在语义-残差正交子空间中表现出低秩坍缩(即秩退化),同时保留主要语义方向。这种结构扁平化在最终解码阶段一致出现,构成跨多种生成架构的共享瓶颈。基于此特征,我们提出LoRC框架,通过解耦语义主导性以捕捉由生成解码瓶颈引起的坍缩残差几何。该方法在多个基准上平均提升7.0%准确率,在39个未见生成器上实现97.0%准确率,展现出强大的跨模型泛化能力与鲁棒性,适用于复杂真实环境中的AIGI检测。
原文摘要 · Abstract (English)
Modern generators faithfully model macroscopic semantics, producing synthetic images that appear highly realistic. Consequently, decisive forensic cues reside in subtle non-semantic visual discrepancies. To reveal these cues, we revisit AIGI detection from a geometric perspective and identify an architecture-agnostic signature. Specifically, modern generators exhibit low-rank collapse (\textit{i.e.}, rank degeneracy) in the semantic-residual orthogonal subspace while largely preserving the dominant semantic direction. This structural flattening consistently emerges during the final decoding stage, forming a shared bottleneck across diverse generator architectures. Motivated by this signature, we propose \textbf{LoRC}, a framework that decouples semantic dominance to capture the collapsed residual geometry induced by the generative decoding bottleneck. Our method improves accuracy by an average of 7.0\% across multiple benchmarks and achieves 97.0\% accuracy on 39 unseen generators. These results demonstrate strong cross-model generalization and robustness, making LoRC a reliable approach for AIGI detection in complex real-world environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。