arXiv:2606.04754cs.LG2026-06中稿 · ICML

揭示神经元可识别性如何让非对称网络仍具线性连接路径

Beyond Structural Symmetries: Linear Mode Connectivity via Neuron Identifiability

论文配图:Beyond Structural Symmetries: Linear Mode Connectivity via Neuron Identifiability
图 1 · 摘自论文原文
  • 基于神经元有效函数类构建理论框架
  • 发现即使结构不对称也能存在大量近似等效解
  • 实现无需对齐的表征融合,且支持低损耗直线路径

深度学习中的诸多现象,如线性模式连通性和训练动态的结构性,与参数对称性密切相关——即变换参数但不改变模型输出的功能。尽管对称性受到越来越多关注,但参数、数据与表征之间的精确关系仍不清晰。为此,我们提出有效函数类的理论框架,即神经元在其输入支持下可实现的函数集合及其对应的范数代价。我们进一步形式化了通过独立训练中神经元可识别性实现的有效对称性破缺。分析表明,即使在结构不对称的模型中,神经网络仍可存在大量近似等效解。此外,我们证明神经元可识别性可实现无需预对齐的表征融合,并刻画了此类融合何时能产生线性低损失路径。这些发现凸显了有效函数类在影响损失曲面中的关键作用。

原文摘要 · Abstract (English)

Many striking phenomena in deep learning, such as linear mode connectivity and the structured behavior of training dynamics, are closely tied to parameter symmetries: transformations that leave the realized function unchanged. Despite growing attention to parameter symmetries, the exact interplay between parameters, data, and representations remains underexplored. To investigate this, we develop a theoretical framework of effective function classes, i.e., the set of functions a neuron can realize on its input support, and the norm cost of realizing them. We then formalize effective symmetry breaking via neuron identifiability across independent training runs. Our analysis shows that neural networks can admit large families of approximately equivalent solutions even in structurally asymmetric models. We further show that neuron identifiability enables representation merging without prior alignment, and characterize when such merging admits a linear low-loss path. These findings highlight the role of effective function classes in affecting the loss landscape.

神经网络损失曲面对称性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。