arXiv:2410.15235cs.CV2024-10被引 2

用自编码器分析图片记忆度,发现可预测记忆效果的视觉特征。

Modeling Visual Memorability Assessment with Autoencoders Reveals Characteristics of Memorable Images

  • 基于VGG16的自编码器学习图像隐含表示,模拟单次曝光记忆实验。
  • 重建误差与记忆度显著相关,隐空间表征独特性也影响记忆效果。
  • 通过梯度分析识别出关键视觉特征,适合关注视觉记忆建模的研究者。

图像记忆度指某些图像比其他图像更易被记住的现象,是可通过单次暴露后回忆概率量化的基本图像属性。尽管人类视觉感知与记忆研究取得进展,但决定图像记忆度的特征仍不明确。为此,本文提出一种基于深度学习的计算建模方法:采用基于VGG16卷积神经网络的自编码器,在单周期训练中学习图像的隐含表示,模拟人类记忆实验中单次暴露后的回忆评估。我们考察了自编码器重建误差与记忆度的关系,分析了隐空间表示的独特性,并构建了多层感知机(MLP)模型进行记忆度预测。此外,利用积分梯度(IG)进行可解释性分析,识别出对记忆度有贡献的关键视觉特征。结果表明,图像记忆度评分与自编码器重建误差存在显著相关性,其隐空间表示具有强预测能力。表征独特性与记忆度显著相关。这些发现表明,自编码器表征捕捉了图像记忆度的核心机制,为人类视觉记忆的计算建模提供了新见解。

原文摘要 · Abstract (English)

Image memorability refers to the phenomenon where certain images are more likely to be remembered than others. It is a quantifiable and intrinsic image attribute, defined as the likelihood of an image being remembered upon a single exposure. Despite advances in understanding human visual perception and memory, it is unclear what features contribute to an image's memorability. To address this question, we propose a deep learning-based computational modeling approach. We employ an autoencoder-based approach built on VGG16 convolutional neural networks (CNNs) to learn latent representations of images. The model is trained in a single-epoch setting, mirroring human memory experiments that assess recall after a single exposure. We examine the relationship between autoencoder reconstruction error and memorability, analyze the distinctiveness of latent space representations, and develop a multi-layer perceptron (MLP) model for memorability prediction. Additionally, we perform interpretability analysis using Integrated Gradients (IG) to identify the key visual characteristics that contribute to memorability. Our results demonstrate a significant correlation between the images' memorability score and the autoencoder's reconstruction error, as well as the robust predictive performance of its latent representations. Distinctiveness in these representations correlated significantly with memorability. Additionally, certain visual characteristics were identified as features contributing to image memorability in our model. These findings suggest that autoencoder-based representations capture fundamental aspects of image memorability, providing new insights into the computational modeling of human visual memory.

图像记忆自编码器可解释性视觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。