视觉与语言模型在深层共享语义表示,且与人类判断一致。
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models
- 通过分析中间层发现两模型在深层产生语义对齐
- 语义变化导致对齐崩溃,证明是真正语义编码
- 多图多文匹配下仍符合人类偏好,适合跨模态研究
近期研究表明,仅训练于单一模态的视觉和语言模型虽未共同训练,却仍会将输入投影到部分对齐的表征空间。然而我们尚不清楚这种对齐在各网络中何时出现、由何种视觉或语言线索支持、是否反映人类在多对多图像-文本场景中的偏好,以及同一概念的多个样本聚合如何影响对齐。本文系统性地探究这些问题。结果表明,对齐在两类模型的中后期层达到峰值,反映了从模态特异性到概念共享表征的转变。该对齐对仅外观的变化鲁棒,但当语义被改变(如物体移除或词序打乱)时即崩溃,说明共享代码确实是语义性的。超越一对一图像-标题范式,强制选择任务‘Pick-a-Pic’显示,人类对图像-标题匹配的偏好在所有视觉-语言模型对的嵌入空间中均被映射。该模式在多个标题对应单个图像时亦双向成立,表明模型能捕捉精细语义差异,接近人类判断。令人惊讶的是,对多个样本的嵌入进行平均反而增强了对齐,而非模糊细节。总体而言,单模态网络收敛于一个与人类判断一致的共享语义码,并随样本聚合而强化。
原文摘要 · Abstract (English)
Recent studies show that deep vision-only and language-only models--trained on disjoint modalities--nonetheless project their inputs into a partially aligned representational space. Yet we still lack a clear picture of where in each network this convergence emerges, what visual or linguistic cues support it, whether it captures human preferences in many-to-many image-text scenarios, and how aggregating exemplars of the same concept affects alignment. Here, we systematically investigate these questions. We find that alignment peaks in mid-to-late layers of both model types, reflecting a shift from modality-specific to conceptually shared representations. This alignment is robust to appearance-only changes but collapses when semantics are altered (e.g., object removal or word-order scrambling), highlighting that the shared code is truly semantic. Moving beyond the one-to-one image-caption paradigm, a forced-choice "Pick-a-Pic" task shows that human preferences for image-caption matches are mirrored in the embedding spaces across all vision-language model pairs. This pattern holds bidirectionally when multiple captions correspond to a single image, demonstrating that models capture fine-grained semantic distinctions akin to human judgments. Surprisingly, averaging embeddings across exemplars amplifies alignment rather than blurring detail. Together, our results demonstrate that unimodal networks converge on a shared semantic code that aligns with human judgments and strengthens with exemplar aggregation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。