视觉与语言的相似性判断共享同一表征空间,反映世界关系的稳定结构。
Representations in vision and language converge in a shared, multidimensional space of perceived similarities
- 用人类行为实验验证视觉与语言相似性在共享空间中收敛
- 行为判断与脑区响应高度一致,且可被LLM嵌入模型解释
- 揭示概念表征源于世界关系而非模态特异性输入
人类能轻松描述所见,但建立视觉与语言间的共享表征仍具挑战。新证据表明,人脑在视觉与语言中的表征可被大型语言模型(LLMs)的语义特征空间良好预测。本研究通过63名参与者对100张自然场景图像及对应句子描述进行相似性判断,发现视觉与语言的相似性判断在行为层面高度收敛,并能预测相同的fMRI脑区响应模式。此外,训练将图像映射至LLM嵌入的计算模型,在解释行为相似性结构上优于类别训练模型和AlexNet对照组。结果表明,人类视觉与语言的相似性判断基于一种跨模态、非模态依赖的共享表征结构,反映了外部世界的稳定关系属性。该发现暗示感官系统与人工系统在概念形成上具有共同机制。
原文摘要 · Abstract (English)
Humans can effortlessly describe what they see, yet establishing a shared representational format between vision and language remains a significant challenge. Emerging evidence suggests that human brain representations in both vision and language are well predicted by semantic feature spaces obtained from large language models (LLMs). This raises the possibility that sensory systems converge in their inherent ability to transform their inputs onto shared, embedding-like representational space. However, it remains unclear how such a space manifests in human behaviour. To investigate this, sixty-three participants performed behavioural similarity judgements separately on 100 natural scene images and 100 corresponding sentence captions from the Natural Scenes Dataset. We found that visual and linguistic similarity judgements not only converge at the behavioural level but also predict a remarkably similar network of fMRI brain responses evoked by viewing the natural scene images. Furthermore, computational models trained to map images onto LLM-embeddings outperformed both category-trained and AlexNet controls in explaining the behavioural similarity structure. These findings demonstrate that human visual and linguistic similarity judgements are grounded in a shared, modality-agnostic representational structure that mirrors how the visual system encodes experience. The convergence between sensory and artificial systems suggests a common capacity of how conceptual representations are formed-not as arbitrary products of first order, modality-specific input, but as structured representations that reflect the stable, relational properties of the external world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。