arXiv:2608.25609cs.CV2026-08

CLIP模型中人类与AI绘画自动分离,源于低感知显著性的多尺度结构差异。

On the Separation of Human and AI-Generated Images in CLIP Embedding Space

论文配图:On the Separation of Human and AI-Generated Images in CLIP Embedding Space
图 1 · 摘自论文原文
  • 通过梯度反演和可解释图像表示,探究嵌入空间中分离机制。
  • 人类难以察觉的微小扰动即可引发显著分离,说明敏感度极高。
  • 揭示了人工视觉与人类感知间存在深层审美差异,适合研究AI生成内容检测者。

我们发现了一个此前未被报道的现象:在联合嵌入空间中,人类与AI生成的绘画自发地沿主导主成分方向分离,且无需专门设计的监督目标。我们的目标并非利用该现象进行检测,而是解释其成因——识别支撑分离的视觉信息,并将其从嵌入空间追溯到图像域。为此,我们结合可解释图像表示与基于梯度的反演方法,系统性地探测特征空间中的关系。鲁棒性实验和越来越复杂的统计描述逐步排除了基于全局图像属性或简单局部统计的直观解释,转而指向分布式多尺度图像结构。多尺度散射提供了最有效的可解释表示,但仍仅部分解释该现象。直接反演揭示了另一关键观察:沿主导CLIP方向产生显著位移的图像扰动,对人类几乎不可察觉,表明这些分离方向对极低感知显著性的图像变化极为敏感。综合结果揭示了CLIP表征所反映的视觉证据与人类感知之间的显著差异,引发关于人工视觉与人类视觉、乃至人工与人类审美判断之间关系的更广泛问题。

原文摘要 · Abstract (English)

We identify a previously unreported phenomenon in CLIP representations: human and AI-generated paintings spontaneously separate along the dominant principal directions of their joint embedding distribution, without any supervised objective designed to distinguish the two classes. Rather than exploiting this phenomenon for detection, our objective is to interpret it: we seek to identify the visual information underlying the separation and to trace it back from the embedding space to the image domain. We pursue this objective through a progressive investigation combining interpretable image representations with gradient-based inversion, used systematically as an experimental probe of the relationships identified in feature space. Robustness experiments and increasingly expressive statistical descriptors progressively rule out several intuitive explanations based on global image properties and simple local statistics, and point instead to distributed multiscale image structure. Multiscale scattering provides the most informative interpretable representation considered, but offers only a partial account of the phenomenon. Direct inversion provides a complementary and striking observation: substantial displacements along the dominant CLIP directions can be induced by image perturbations that remain nearly imperceptible to human observers, showing that the directions involved in the separation are highly sensitive to image variations with very low perceptual salience for humans. Taken together, these results reveal a significant difference between the visual evidence reflected in CLIP representations and that readily accessible to human perception, raising broader questions about the relationship between artificial and human vision and, ultimately, between artificial and human aesthetic judgment.

CLIP图像生成美学差异可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。