扩散模型能模拟人类视觉错觉,还能生成新错觉图像。
The Art of Deception: Color Visual Illusions and Diffusion Models
- 在扩散模型隐空间中发现类人亮度/颜色错觉
- 模型可准确预测已有视觉错觉,且生成新错觉
- 生成的错觉能骗过真实人类,验证其真实性
人类视觉错觉源于对分布外刺激的误判:当观察者适应特定统计规律时,异常刺激的感知会偏离现实。近期研究发现人工神经网络(ANNs)也会被视觉错觉欺骗。这引发深层问题:为何人类大脑与ANN都易受相同错觉影响?任何ANN是否都应具备错觉感知能力?这种感知是优势还是缺陷?本文研究视觉错觉在扩散模型中的编码方式。令人惊讶的是,扩散模型在隐空间中表现出类人的亮度与颜色偏移。利用这一特性,我们证明扩散模型可预测视觉错觉,并通过文本到图像扩散模型生成前所未见的真实感错觉图像。通过心理物理学实验验证,模型生成的错觉同样能欺骗人类观察者。
原文摘要 · Abstract (English)
Visual illusions in humans arise when interpreting out-of-distribution stimuli: if the observer is adapted to certain statistics, perception of outliers deviates from reality. Recent studies have shown that artificial neural networks (ANNs) can also be deceived by visual illusions. This revelation raises profound questions about the nature of visual information. Why are two independent systems, both human brains and ANNs, susceptible to the same illusions? Should any ANN be capable of perceiving visual illusions? Are these perceptions a feature or a flaw? In this work, we study how visual illusions are encoded in diffusion models. Remarkably, we show that they present human-like brightness/color shifts in their latent space. We use this fact to demonstrate that diffusion models can predict visual illusions. Furthermore, we also show how to generate new unseen visual illusions in realistic images using text-to-image diffusion models. We validate this ability through psychophysical experiments that show how our model-generated illusions also fool humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。