语言是多模态表示收敛的终极吸引点,非语言模态总向语言靠拢。
The Wittgensteinian Representation Hypothesis: Is Language the Attractor of Multimodal Convergence?

- 用循环kNN检测方向性,发现非语言模态明显向语言表示靠近
- 无论模型规模或类型,语言始终处于表征空间最紧凑区域
- 揭示语言结构本质为压缩优化下的自然吸引子,适合关注多模态对齐的研究者
理解不同模态的独立训练神经网络为何收敛到共享表示,以及收敛目标是什么,仍是表示学习中的未解之谜。现有研究依赖对称相似性度量,虽能检测收敛却无法判断方向。本文引入基于循环kNN的非对称对齐度量,分析数十个独立训练的单模态模型(涵盖点云、视觉、语言)。结果揭示一致的方向不对称性:非语言模态显著更倾向于靠近语言的邻域结构,而反向则不成立;该现象在所有模型族和尺度下均成立,但对称度量完全无法捕捉。机制分析表明方向性源于特征密度不对称——语言表示占据表征空间中最紧凑的区域。信息瓶颈框架提供理论解释:压缩约束下的优化会驱动表示趋向离散、组合性的语言特征。我们形式化提出「维特根斯坦表示假设」:语言的语义结构是多模态表示收敛的渐近吸引子。
原文摘要 · Abstract (English)
Understanding why independently trained neural networks from different modalities converge toward shared representations, and where this convergence leads, remains an open question in representation learning. All existing evidence relies on symmetric similarity measures, which can detect convergence but are structurally blind to its direction. We introduce directional convergence analysis using cycle-kNN, an asymmetric alignment measure, applied across dozens of independently trained unimodal models spanning point clouds, vision, and language. We uncover a consistent directional asymmetry: non-language modalities move toward the neighborhood structure of language significantly more than the reverse, and this pattern holds across all model families and scales--yet is entirely invisible to symmetric measures. Mechanistic analysis traces the directionality to feature density asymmetry, whereby language representations occupy the most compact regions of representational space. The Information Bottleneck framework provides a principled interpretation: optimization under compression drives representations toward discrete, compositional structures characteristic of language. We formalize this as the Wittgensteinian Representation Hypothesis: the semantic structure of language is the asymptotic attractor of multimodal representation convergence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。