arXiv:2409.02347cs.LG2024-09

通过新方法探索权重集成中功能多样性的作用,提升模型泛化能力。

Understanding the Role of Functional Diversity in Weight-Ensembling with Ingredient Selection and Multidimensional Scaling

  • 提出两种新集成方法,基于功能多样性选择模型组件。
  • 实验证明高多样性显著提升集成性能,但非唯一决定因素。
  • 适合关注模型集成与泛化能力的研究者参考。

权重集成通过直接平均多个神经网络的参数形成单一模型,表现出良好的分布内(ID)和分布外(OOD)泛化能力,但其原理尚未完全理解。尽管普遍认为成功得益于各模型提供的功能多样性,但如何选取最优组合仍不明确,当前最优方法为线性时间的‘贪心’策略。本文引入两种新型权重集成方法,研究性能动态与算法如何利用功能多样组件之间的关系,类似于预测集成文献中的多样性鼓励机制。开发可视化工具,通过成对距离定义的领域探索,分析算法的选择行为与收敛特性。实证分析揭示:高多样性有助于提升权重集成效果,但需结合位置差异性模型选择才能实现最佳性能。同时证明,采样位置相异的模型同样能有效提升集成表现。

原文摘要 · Abstract (English)

Weight-ensembles are formed when the parameters of multiple neural networks are directly averaged into a single model. They have demonstrated generalization capability in-distribution (ID) and out-of-distribution (OOD) which is not completely understood, though they are thought to successfully exploit functional diversity allotted by each distinct model. Given a collection of models, it is also unclear which combination leads to the optimal weight-ensemble; the SOTA is a linear-time ``greedy" method. We introduce two novel weight-ensembling approaches to study the link between performance dynamics and the nature of how each method decides to use apply the functionally diverse components, akin to diversity-encouragement in the prediction-ensemble literature. We develop a visualization tool to explain how each algorithm explores various domains defined via pairwise-distances to further investigate selection and algorithms' convergence. Empirical analyses shed perspectives which reinforce how high-diversity enhances weight-ensembling while qualifying the extent to which diversity alone improves accuracy. We also demonstrate that sampling positionally distinct models can contribute just as meaningfully to improvements in a weight-ensemble.

权重集成功能多样性模型泛化集成学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。