arXiv:2502.17391cs.LGcs.AI2025-02中稿 · ICLR被引 1

让神经网络打破对称性,能显著提升集成模型性能。

The Empirical Impact of Reducing Symmetries on the Performance of Deep Ensembles and MoE

  • 通过打破网络对称性,提升深度集成模型表现
  • 对称性越弱,集成规模越大时性能越优
  • 适合关注集成学习与模型泛化的研究者

近期研究表明,减少神经网络中的对称性可增强网络间的线性模式连通性,且无需参数空间对齐,从而提升线性插值网络的性能。然而在实际应用中,神经网络插值很少使用,而网络集成更为常见。本文在五个数据集上实证研究了减少对称性对深度集成和专家混合(MoE)模型的影响。为进一步探索线性模式连通性,我们提出混合插值专家(MoIE)架构。结果表明,基于非对称神经网络构建的深度集成,在集成规模增大时性能显著优于对称版本;但关于对称性减少是否影响MoE与MoIE架构,实验未提供明确结论。

原文摘要 · Abstract (English)

Recent studies have shown that reducing symmetries in neural networks enhances linear mode connectivity between networks without requiring parameter space alignment, leading to improved performance in linearly interpolated neural networks. However, in practical applications, neural network interpolation is rarely used; instead, ensembles of networks are more common. In this paper, we empirically investigate the impact of reducing symmetries on the performance of deep ensembles and Mixture of Experts (MoE) across five datasets. Additionally, to explore deeper linear mode connectivity, we introduce the Mixture of Interpolated Experts (MoIE). Our results show that deep ensembles built on asymmetric neural networks achieve significantly better performance as ensemble size increases compared to their symmetric counterparts. In contrast, our experiments do not provide conclusive evidence on whether reducing symmetries affects both MoE and MoIE architectures.

深度集成对称性MoE模型性能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。