揭示神经网络在训练与分布变化中表征对齐的演化规律,发现早期层对齐稳定,深层随数据漂移发散。
Bridging Critical Gaps in Convergent Learning: How Representational Alignment Evolves Across Layers, Training, and Distribution Shifts
- 对比三类对齐方法,发现正交变换已足够有效,排列匹配显著优于随机水平。
- 90%以上表征对齐在首轮训练内完成,远早于准确率饱和,受输入统计和结构偏差驱动。
- 分布外测试中浅层保持强对齐,深层对齐随数据偏移程度线性下降,适用于跨模型分析。
理解收敛学习——即独立训练的神经网络系统(如多个深度模型或脑与模型)是否形成相似内部表示——对神经科学与人工智能均至关重要。然而现有研究范围狭窄:通常仅比较少数模型、单一数据集,依赖一种对齐度量,并仅在单一时点评估网络。本文开展大规模审计,涵盖数十个视觉模型与数千对层间比较,填补长期空白。首先,我们对比三类对齐方法:线性回归(仿射不变)、正交普鲁克斯特(旋转/反射不变)与排列/软匹配(单位顺序不变)。结果表明,正交变换的对齐效果几乎等同于更灵活的线性方法;尽管排列得分较低,但仍显著高于随机水平,表明存在特定的表征基。追踪训练过程发现,几乎所有最终对齐均在第一轮训练内完成,远早于准确率趋于平稳,说明其主要由共享输入统计与架构偏差驱动,而非最终任务解。最后,在多种分布外图像测试下,浅层仍保持紧密对齐,而深层对齐程度随分布偏移量线性下降。这些发现深化了对表征收敛机制的理解,对神经科学与人工智能均有启示。
原文摘要 · Abstract (English)
Understanding convergent learning -- the degree to which independently trained neural systems -- whether multiple artificial networks or brains and models -- arrive at similar internal representations -- is crucial for both neuroscience and AI. Yet, the literature remains narrow in scope -- typically examining just a handful of models with one dataset, relying on one alignment metric, and evaluating networks at a single post-training checkpoint. We present a large-scale audit of convergent learning, spanning dozens of vision models and thousands of layer-pair comparisons, to close these long-standing gaps. First, we pit three alignment families against one another -- linear regression (affine-invariant), orthogonal Procrustes (rotation-/reflection-invariant), and permutation/soft-matching (unit-order-invariant). We find that orthogonal transformations align representations nearly as effectively as more flexible linear ones, and although permutation scores are lower, they significantly exceed chance, indicating a privileged representational basis. Tracking convergence throughout training further shows that nearly all eventual alignment crystallizes within the first epoch -- well before accuracy plateaus -- indicating it is largely driven by shared input statistics and architectural biases, not by the final task solution. Finally, when models are challenged with a battery of out-of-distribution images, early layers remain tightly aligned, whereas deeper layers diverge in proportion to the distribution shift. These findings fill critical gaps in our understanding of representational convergence, with implications for neuroscience and AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。