arXiv:2607.17138cs.CV2026-07

不同架构的去噪模型内部会形成类似人类的错觉感知表征,但这些表征不体现在输出上。

Denoising Models Develop Human-Like Perceptual Illusion Representations Across Architectures

论文配图:Denoising Models Develop Human-Like Perceptual Illusion Representations Across Architectures
图 1 · 摘自论文原文
  • 在特定内部层识别出对错觉敏感的神经元通道
  • 内部激活与人类亮度感知模型相关性达ρ≥0.70,且随错觉强度单调变化
  • 这类表征是内部处理的关键,但无法通过输出检测,称作感知幻象

在自然图像上训练的深度神经网络在亮度错觉任务中表现出与人类观察者一致的输出。尽管这一现象已在多种架构中被验证,但所有证据均来自输出层面:恢复像素、解码轨迹或分类决策。模型是否在内部真正表征错觉,以及具体在何处、如何实现,仍不清楚。我们发现,不同架构的去噪模型在特定内部层发展出对错觉敏感的表征,能区分错觉区域与物理匹配的对照区域。去噪目标比架构更关键。在适配刺激下,这些激活与已验证的人类亮度感知模型(FLODOG)高度相关(斯皮尔曼ρ≥0.70),且随错觉强度单调增长。通过通道消融实验,我们提供因果证据,表明这些通道显著影响内部信号。然而,将这些表征注入生成流程后,所有测试架构均未产生可测量的像素变化;我们将其称为‘感知幻象’:在内部处理中活跃,却无法通过输出评估察觉。此前此类内部-输出分离现象仅见于语言模型,这是首次在去噪视觉模型的感知表征中发现。

原文摘要 · Abstract (English)

Deep neural networks trained on natural images are shown to produce outputs consistent with human observers for brightness illusions. While this phenomenon has been documented across architectures, all evidence, to date, is measured at the output level: restored pixels, decoded trajectories, or classification decisions. Whether these models actually represent illusions internally, and if so where and how, remains unknown. We show that denoising models develop illusion-sensitive representations at specific internal layers, across varied architectures. Specifically, we identify the layers and channels that discriminate illusory from physically matched control regions. We show that the denoising objective is a more important driver of the effect than the architecture. On domain-appropriate stimuli, these activations track a validated psychophysical model of human brightness perception (FLODOG; Spearman $ρ\geq 0.70$) and scale monotonically with parametric illusion strength. Leveraging these findings, we provide causal evidence via channel ablation showing that illusion-sensitive channels specifically and substantially affect the internal signal. Yet injecting these representations into the generation pipeline produces no measurable pixel shift across all tested architectures; we term such representations perceptual phantoms: active in internal processing yet invisible to any output-based evaluation. While related internal-output dissociations have been characterized in language models, this is the first such characterization for perceptual representations in denoising vision models.

去噪模型感知错觉内部表征感知幻象

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。