对比大模型与视觉语言模型在阅读时的脑活动契合度,发现后者更贴近人类行为。
Do VLMs Align Better with Humans than LLMs during Natural Reading?

- 用相同文本输入比较大模型与视觉语言模型的阅读表现。
- 视觉语言模型更准确预测人类回视眼动,尤其在画面感强的句子中。
- 该优势源于视觉语义训练,适合关注多模态与认知对齐的研究者。
大型语言模型在模拟人类语言处理方面日益重要,但视觉-语言训练是否使文本表征更接近人类自然阅读仍不明确。我们通过在纯文本输入下对比匹配的LLM与视觉语言模型(VLM)对人脑活动(全皮层fMRI)和行为(同步回视眼动)的拟合程度,研究这一问题。结果表明,视觉语言训练对模型-人类对齐具有选择性而非全局性影响:在两组同源模型中,VLMs更能准确预测人类回视眼动,而全皮层fMRI对齐度则相近;但句级分析显示,当句子视觉意象越强时,VLM在fMRI对齐上的优势越大。这些发现提供了多模态训练历史的受控计算比较,表明视觉语言预训练能通过阅读行为和具象内容提升模型与人类的对齐效果。
原文摘要 · Abstract (English)
Large language models have become increasingly useful computational models of human language processing, but it remains open whether vision-language learning makes text representations more human-like during natural reading. We address this question by comparing matched LLM and vision-language model pairs under strictly text-only input and evaluating alignment with human brain activity (whole-cortex fMRI) and human behavior (synchronized regressive saccades). We identify a selective, rather than global, effect of vision-language training on human-model alignment. In the two within-lineage model pairs, VLMs more accurately predicted human regressive saccades, whereas VLMs and LLMs showed comparable whole-cortex fMRI alignment. However, sentence-level analyses revealed that the VLM advantage in fMRI alignment increased with the visual evocative strength of the sentences. Together, these findings provide a controlled in-silico comparison of multimodal training histories, showing that vision-language pretraining selectively improves model-human alignment via reading behavior and visually grounded content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。