揭示Transformer模型间低损耗连接路径,突破传统对称性限制
Generalized Linear Mode Connectivity for Transformers
- 提出四类对称变换统一框架,涵盖参数空间更广的重参数化方法
- 首次实现独立训练的ViT与GPT-2间零损耗线性路径连接
- 适用于不同架构、宽度的多模型对齐,推动损失曲面几何理解
理解神经网络损失曲面的几何结构是深度学习的核心问题,影响泛化与优化。一个显著现象是线性模式连通性(LMC),即独立训练的模型可通过低损或零损路径相连,尽管看似位于不同的损失盆地。然而,参数空间中的对称性(如神经元排列)常使功能等价的模型显得不同。以往工作主要关注通过排列进行神经元重排,但其范围有限,无法捕捉Transformer等现代架构的丰富对称性。本文提出统一框架,涵盖四类对称性:排列、半排列、正交变换及一般可逆映射,扩展了有效重参数化集合,并包含多数已有方法作为特例。关键突破在于首次发现独立训练的Vision Transformer与GPT-2之间存在低损甚至零损的线性插值路径。该框架还可推广至多模型及宽窄异构设置,实现跨不同规模架构的对齐。结果揭示了损失曲面更深层结构,强调了对称性感知分析在理解模型空间几何中的重要性。
原文摘要 · Abstract (English)
Understanding the geometry of neural network loss landscapes is a central question in deep learning, with implications for generalization and optimization. A striking phenomenon is linear mode connectivity (LMC), where independently trained models can be connected by low- or zero-loss paths despite appearing to lie in separate loss basins. However, this is often obscured by symmetries in parameter space -- such as neuron permutations -- which make functionally equivalent models appear dissimilar. Prior work has predominantly focused on neuron reordering through permutations, but such approaches are limited in scope and fail to capture the richer symmetries exhibited by modern architectures such as Transformers. In this work, we introduce a unified framework that captures four symmetry classes -- permutations, semi-permutations, orthogonal transformations, and general invertible maps -- broadening the set of valid reparameterizations and subsuming many previous approaches as special cases. Crucially, this generalization enables, for the first time, the discovery of low- and zero-barrier linear interpolation paths between independently trained Vision Transformers and GPT-2 models. Furthermore, our framework extends beyond pairwise alignment to multi-model and width-heterogeneous settings, enabling alignment across architectures of different sizes. These results reveal deeper structure in the loss landscape and underscore the importance of symmetry-aware analysis for understanding model space geometry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。