用视觉编码器分析图像记忆度,发现重构损失是关键预测指标。
Correlates of Image Memorability in Vision Encoders: Activations, Attention Entropy, Patch Uniformity and Autoencoder Losses
- 通过激活值、注意力熵和补丁均匀性探索模型内部特征
- 自编码器重构损失比传统方法更准确预测图像记忆度
- 适合研究视觉记忆机制或模型可解释性的研究人员
图像在人类记忆中的留存程度各不相同。受认知科学与计算机视觉研究启发,我们首次系统探究预训练的基于Transformer的视觉编码器中与图像记忆度相关的内部特征。重点关注激活值、注意力分布及图像补丁的均匀性,发现这些特征与记忆度存在一定相关性。此外,我们引入稀疏自编码器对视觉编码器表示进行建模,其重构损失作为记忆度的代理指标,表现优于以往基于卷积神经网络的方法。结果表明,某些内部特征可有效预测图像的人类记忆度,尤其自编码器的重建损失具有强相关性。
原文摘要 · Abstract (English)
Images vary in how memorable they are to humans. Inspired by findings from cognitive science and computer vision, we explore correlates of image memorability in pretrained transformer-based vision encoders for the first time. Focusing initially on activations, attention distributions, and the uniformity of image patches, we find that these features correlate with memorability to some extent. Additionally, we explore sparse autoencoder loss over the representations of vision encoders as a proxy for memorability, which yields results outperforming past methods using convolutional neural network representations. Our results shed light on the relationship between model-internal features and memorability. They show that some features are informative predictors of what makes images memorable to humans; revealing that, in particular, the reconstruction loss from our autoencoders is a strong correlate of image memorability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。