arXiv:2410.23603cs.CVcs.CL2024-10被引 2

用深度模型分离视觉美感与语言描述,发现感知远比语言重要

Using Multimodal Deep Neural Networks to Disentangle Language from Visual Aesthetics

  • 通过多模态神经网络解耦视觉与语言对美感的影响
  • 纯视觉模型解释了90%以上美感评分的方差
  • 语言描述无法提升美感预测,说明美在感知而非表达

当我们感知一幅图像的美时,这种体验中有多少来自难以言说的感知计算,又有多少源于可被自然语言表达的概念知识?通过行为实验或神经影像手段分离感知与语言在审美体验中的作用,往往难以实现。本文采用线性解码方法,基于单模态视觉、单模态语言及多模态(语言对齐)深度神经网络(DNN)模型的特征表示,预测人类对自然图像的美感评分。结果显示,单模态视觉模型(如SimCLR)解释了绝大多数可解释方差;语言对齐视觉模型(如SLIP)仅带来微小提升;以视觉嵌入为条件生成描述的单模态语言模型(如CLIPCap)未进一步增益;仅使用描述嵌入的预测性能低于图像与描述嵌入拼接后的联合表示。这些结果表明,即便我们能用语言描述美感,其核心仍来自前馈感知的不可言说计算。

原文摘要 · Abstract (English)

When we experience a visual stimulus as beautiful, how much of that experience derives from perceptual computations we cannot describe versus conceptual knowledge we can readily translate into natural language? Disentangling perception from language in visually-evoked affective and aesthetic experiences through behavioral paradigms or neuroimaging is often empirically intractable. Here, we circumnavigate this challenge by using linear decoding over the learned representations of unimodal vision, unimodal language, and multimodal (language-aligned) deep neural network (DNN) models to predict human beauty ratings of naturalistic images. We show that unimodal vision models (e.g. SimCLR) account for the vast majority of explainable variance in these ratings. Language-aligned vision models (e.g. SLIP) yield small gains relative to unimodal vision. Unimodal language models (e.g. GPT2) conditioned on visual embeddings to generate captions (via CLIPCap) yield no further gains. Caption embeddings alone yield less accurate predictions than image and caption embeddings combined (concatenated). Taken together, these results suggest that whatever words we may eventually find to describe our experience of beauty, the ineffable computations of feedforward perception may provide sufficient foundation for that experience.

多模态美学计算感知解耦深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。