大规模数据下跨模态表示并不收敛,语言与视觉模型学的是不同世界。
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

- 在百万级数据上测试跨模态对齐,发现小样本结论不可靠
- 多模态对齐仅存粗粒度语义重叠,无精细结构一致性
- 新模型不呈现更强对齐,适合关注多模态机制的研究者
柏拉图表征假说认为,训练于不同模态(如文本与图像)的神经网络会趋向于同一现实表征。若成立,模态选择将无关紧要。我们发现该假说的实验证据极为脆弱,严重依赖评估设置。对齐度在约1千样本的小数据集上较高,但扩展至数百万样本后显著下降。此现象不仅存在于文本-图像,也见于文本-音频和文本-视频对齐。剩余对齐仅反映粗粒度语义重叠,而非一致的细粒度结构。此外,Huh等人的评估采用一对一图像-标题设定,这一限制在真实场景的多对多关系中失效,进一步降低对齐测量值。我们还发现,近年更强的语言模型并未表现出与视觉模型更强的对齐趋势。总体而言,当前跨模态表征收敛的证据远弱于后续研究所认为的。不同模态训练的模型可能学习到同样丰富的世界表征,但并非同一表征。
原文摘要 · Abstract (English)
The Platonic Representation Hypothesis suggests that neural networks trained on different modalities (e.g., text and images) align and eventually converge toward the same representation of reality. If true, this has significant implications for whether modality choice matters at all. We show that the experimental evidence for this hypothesis is fragile and depends critically on the evaluation regime. Alignment is measured using mutual nearest neighbors on small datasets ($\approx$1K samples) and degrades substantially as the dataset is scaled to millions of samples. The same behavior is observed beyond text-image, for text-audio and text-video alignment. The alignment that remains between model representations reflects coarse semantic overlap rather than consistent fine-grained structure. Moreover, the evaluations in Huh et al. are done in a one-to-one image-caption setting, a constraint that breaks down in realistic many-to-many settings and further reduces measured alignment. We also find that the reported trend of stronger language models increasingly aligning with vision does not appear to hold for newer models. Overall, our findings suggest that the current evidence for cross-modal representational convergence is considerably weaker than subsequent works have taken it to be. Models trained on different modalities may learn equally rich representations of the world, just not the same one.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。