用单个刺激的视觉模型一致性预测跨模态对齐程度。
Modulating Cross-Modal Convergence with Single-Stimulus, Intra-Modal Dispersion

- 基于广义普罗克拉斯忒斯算法,量化单个刺激下的视觉表征一致性。
- 视觉模型共识度高的刺激使图文模型对齐度提升最多达两倍。
- 该方法揭示了神经网络与人类脑表征收敛的潜在机制,适合研究跨模态对齐者。
神经网络在不同架构、训练目标甚至数据模态下表现出显著的表征收敛性,这种收敛性可预测其与大脑表征的一致性。近期假说认为这是由于网络以相似方式学习环境内在结构所致。然而,单个刺激如何引发跨网络的收敛表征仍不明确。一张图像可能有多种感知方式,也可用不同语言表达。本文提出一种基于广义普罗克拉斯忒斯算法的方法,在单刺激层面测量视觉模态内的表征收敛性。我们对具有不同训练目标的视觉模型进行测试,选取基于其内部一致性(模态内离散度)的刺激。关键发现:模态内离散度越低(视觉模型间共识越高),视觉与语言模型间的跨模态对齐越高,最高可达两倍差异(如DINOv2与语言模型配对)。该现象在不同刺激选择标准和模型组合中均稳健。在单刺激层面测量收敛性,为理解模态间及神经网络与人脑表征的收敛与分歧提供了新路径。
原文摘要 · Abstract (English)
Neural networks exhibit a remarkable degree of representational convergence across diverse architectures, training objectives, and even data modalities. This convergence is predictive of alignment with brain representation. A recent hypothesis suggests this arises from learning the underlying structure in the environment in similar ways. However, it is unclear how individual stimuli elicit convergent representations across networks. An image can be perceived in multiple ways and expressed differently using words. Here, we introduce a methodology based on the Generalized Procrustes Algorithm to measure intra-modal representational convergence at the single-stimulus level. We applied this to vision models with distinct training objectives, selecting stimuli based on their degree of alignment (intra-modal dispersion). Crucially, we found that this intra-modal dispersion strongly modulates alignment between vision and language models (cross-modal convergence). Specifically, stimuli with low intra-modal dispersion (high agreement among vision models) elicited significantly higher cross-modal alignment than those with high dispersion, by up to a factor of two (e.g., in pairings of DINOv2 with language models). This effect was robust to stimulus selection criteria and generalized across different pairings of vision and language models. Measuring convergence at the single-stimulus level provides a path toward understanding the sources of convergence and divergence across modalities, and between neural networks and human neural representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。