用模型集成生成视觉幻觉图,揭示不同网络的感知不变性差异
Understanding Cross-Model Perceptual Invariances Through Ensemble Metamers
- 用多个模型集成生成在神经激活上相同但外观不同的图像
- 卷积网络生成的幻觉图更像人眼识别的物体,视觉变压器则更真实但难迁移
- 适合研究模型可解释性与人类视觉对齐的学者参考
理解人工神经网络的感知不变性对于提升可解释性及使模型与人类视觉对齐至关重要。幻觉刺激(metamers)——物理上不同但引发相同神经激活的输入——是研究这些不变性的有效工具。本文提出一种新方法,通过集成多个异构人工神经网络(包括卷积神经网络与视觉变压器)生成幻觉图,捕捉跨不同架构的共享表征子空间。为评估生成幻觉图的特性,我们采用一系列基于图像的度量标准,评估语义保真度与自然性。结果表明,卷积神经网络生成的幻觉图更具可识别性和人类感知相似性,而视觉变压器生成的幻觉图虽更真实但迁移性较差,凸显了网络架构偏差对表征不变性的影响。
原文摘要 · Abstract (English)
Understanding the perceptual invariances of artificial neural networks is essential for improving explainability and aligning models with human vision. Metamers - stimuli that are physically distinct yet produce identical neural activations - serve as a valuable tool for investigating these invariances. We introduce a novel approach to metamer generation by leveraging ensembles of artificial neural networks, capturing shared representational subspaces across diverse architectures, including convolutional neural networks and vision transformers. To characterize the properties of the generated metamers, we employ a suite of image-based metrics that assess factors such as semantic fidelity and naturalness. Our findings show that convolutional neural networks generate more recognizable and human-like metamers, while vision transformers produce realistic but less transferable metamers, highlighting the impact of architectural biases on representational invariances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。