arXiv:2503.13212cs.LG2025-03中稿 · Journal of Vision

通过人眼反馈直接探索人类视觉的多维相似刺激空间。

MAME: Multidimensional Adaptive Metamer Exploration with Human Perceptual Feedback

  • 基于神经网络响应动态调节图像,结合人眼感知反馈自适应生成
  • 发现低层特征生成的相似图人类更难区分,表明早期视觉处理更关键
  • 为研究人类视觉功能组织提供可直接验证的新工具,适合视觉认知研究者

人类脑网络与人工模型之间的对齐已成为视觉科学与机器学习领域的热点。传统方法依赖生物启发模型间接推断人类相似刺激(metamers),但缺乏直接搜索人类相似空间的手段。本文提出多维自适应相似探索框架MAME,通过在线图像生成并结合人眼感知反馈,实现对人类多维相似空间的直接高维探索。MAME根据分层神经网络响应调节参考图像,并依据参与者感知辨别能力自适应更新生成参数。单次心理物理学实验即成功测量出人类多维相似空间。使用生物合理CNN模型的实验表明,基于低层CNN特征的Gram矩阵生成的相似图像,其人类辨别敏感度低于高层特征生成的图像,说明低层处理中人类与模型的相似空间对齐程度较差。这一反直觉发现强调了早期视觉计算在构建生物合理性模型中的重要性。MAME可作为未来直接研究人类视觉功能组织的科学工具。

原文摘要 · Abstract (English)

Alignment between human brain networks and artificial models has become an active research area in vision science and machine learning. A widely adopted approach is identifying "metamers," stimuli physically different yet perceptually equivalent within a system. However, conventional methods lack a direct approach to searching for the human metameric space. Instead, researchers first develop biologically inspired models and then infer about human metamers indirectly by testing whether model metamers also appear as metamers to humans. Here, we propose the Multidimensional Adaptive Metamer Exploration (MAME) framework, enabling direct, high-dimensional exploration of human metameric spaces through online image generation guided by human perceptual feedback. MAME modulates reference images across multiple dimensions based on hierarchical neural network responses, adaptively updating generation parameters according to participants' perceptual discriminability. Using MAME, we successfully measured multidimensional human metameric spaces within a single psychophysical experiment. Experimental results using a biologically plausible CNN model showed that human discrimination sensitivity was lower for metameric images based on Gram-matrix representations derived from low-level CNN features than for those derived from high-level CNN features. The finding suggests a relatively worse alignment between the metameric spaces of humans and the CNN model for low-level processing compared to high-level processing. Counterintuitively, given recent discussions on alignment at higher representational levels, our results highlight the importance of early visual computations in shaping biologically plausible models. Our MAME framework can serve as a future scientific tool for directly investigating the functional organization of human vision.

视觉感知相似刺激神经网络人机对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。