arXiv:2601.11675cs.CVcs.AI2026-01ICLR

用视觉焦点与周边信息生成人眼理解的图像幻影,揭示人类场景认知机制。

Generating metamers of human scene understanding

  • 基于双流特征融合,结合注视点高分辨与周边低分辨信息生成图像。
  • 在行为实验中,92%的生成图被判断为与原图相同,验证了其感知一致性。
  • 适用于研究视觉认知、神经科学及人机交互中的场景理解模型。

人类视觉将视网膜周边区域的低分辨率‘整体感’信息与注视位置的稀疏高分辨率信息相结合,形成对视觉场景的整体理解。本文提出MetamerGen,一种生成与人类潜在场景表征一致的图像幻影的工具。MetamerGen是一种潜空间扩散模型,通过融合周边获取的场景整体信息与注视区域的信息,生成人类观看后所理解的场景图像幻影。从高低分辨率(即‘中心聚焦’)输入生成图像构成一项新颖的图像到图像合成任务,我们通过引入双流表示来解决:使用DINOv2 tokens融合注视区的详细特征与周边退化的上下文特征。为评估生成图像与潜在人类场景表征之间的感知对齐性,我们进行了‘相同-不同’行为实验,参与者需判断生成图与原图是否相同。实验结果显示,生成图像在感知上与原始场景高度一致,识别出真正符合观者潜意识场景表征的幻影。该方法不仅可生成随机注视下的幻影,但当生成条件基于观者自身注视区域时,高层语义对齐性最强,最能预测幻影效果。本研究为理解场景认知提供了有力工具,初步分析揭示了多层级视觉处理中影响人类判断的关键特征。

原文摘要 · Abstract (English)

Human vision combines low-resolution "gist" information from the visual periphery with sparse but high-resolution information from fixated locations to construct a coherent understanding of a visual scene. In this paper, we introduce MetamerGen, a tool for generating scenes that are aligned with latent human scene representations. MetamerGen is a latent diffusion model that combines peripherally obtained scene gist information with information obtained from scene-viewing fixations to generate image metamers for what humans understand after viewing a scene. Generating images from both high and low resolution (i.e. "foveated") inputs constitutes a novel image-to-image synthesis problem, which we tackle by introducing a dual-stream representation of the foveated scenes consisting of DINOv2 tokens that fuse detailed features from fixated areas with peripherally degraded features capturing scene context. To evaluate the perceptual alignment of MetamerGen generated images to latent human scene representations, we conducted a same-different behavioral experiment where participants were asked for a "same" or "different" response between the generated and the original image. With that, we identify scene generations that are indeed metamers for the latent scene representations formed by the viewers. MetamerGen is a powerful tool for understanding scene understanding. Our proof-of-concept analyses uncovered specific features at multiple levels of visual processing that contributed to human judgments. While it can generate metamers even conditioned on random fixations, we find that high-level semantic alignment most strongly predicts metamerism when the generated scenes are conditioned on viewers' own fixated regions.

场景理解图像生成认知模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。