用语言模型嵌入图像描述,能更好预测人脑视觉反应。
What Makes Linguistic Representations Good Models of High-Level Visual Perception in the Human Brain?

- 用不同来源的图像描述和五种语言模型对比测试
- 文本嵌入模型比自回归模型预测人脑反应更准
- 中间层表示最接近人脑与行为判断,适合研究高级视觉
使用语言模型(LMs)对图像描述进行编码,可有效预测人脑在高级视觉区域对自然图像的响应,但影响预测能力的关键因素尚不明确。本研究系统分析了六种不同类型的图像描述(包括人工标注和多种机器生成描述),并用五种语言模型进行表征。结果发现,机器生成的描述在预测人脑活动和行为相似性判断方面表现优异,常优于以往研究中使用的人工标注描述。无论在脑活动预测还是行为对齐上,文本嵌入模型(即微调过的句子级表征模型)均显著优于自回归模型。进一步分析表明,模型中间层的表示在预测能力和行为一致性上达到峰值,该位置恰在句法与语义结构出现之后。结果说明,图像描述的内容及其语言模型表征方式共同影响建模效果,证实了使用描述嵌入研究高级视觉感知的有效性。
原文摘要 · Abstract (English)
Image descriptions represented with language models (LMs) predict human brain responses to naturalistic images in high-level visual regions, but the factors driving this predictivity remain unclear. To investigate this, we systematically studied how images are described and which language models are used to embed those descriptions. For a common set of images, we considered six caption types -- including human-annotated and multiple machine-generated captions -- differing along several dimensions. Each caption was represented with five LMs, spanning autoregressive LMs trained to predict upcoming words and text embedders, i.e., LMs fine-tuned on semantic tasks requiring sentence/document-level representations. Machine-generated captions yielded significant brain predictivity and alignment, often surpassing human-annotated captions used in previous work. Across caption types, text embedders consistently outperformed autoregressive LMs, a pattern replicated when measuring behavioural alignment with image-similarity judgments. Analyses of caption representations from different model layers further revealed that both brain predictivity and behavioural alignment peak at intermediate network depth, shortly after a point thought to mark the emergence of syntactic and semantic structure. Altogether, our results demonstrate that both the content of image captions and the LM used to represent them influence brain- and behaviour-modelling performance, establishing caption embeddings as a useful tool for studying high-level visual perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。