让神经网络权重生成模型更懂对称性,提升效率与泛化能力。
Geometric Flow Models over Neural Network Weights
- 基于流匹配和图神经网络,设计三种尊重权重对称性的生成模型。
- 参数量少一个数量级,仍保持竞争力并可进一步扩展性能。
- 适合做贝叶斯深度学习、迁移学习等任务的权重建模研究者。
深度生成模型如流模型和扩散模型在建模视频、蛋白质等高维复杂数据方面表现优异,这推动了其在神经网络权重等不同数据模态中的应用。权重空间的生成模型对贝叶斯深度学习、学习优化和迁移学习等任务具有重要意义。然而,现有权重空间生成模型常忽略神经网络权重的对称性,或仅考虑部分对称性。建模如MLP层间排列对称性、卷积核滤波器对称性以及非线性激活带来的缩放对称性,有望通过降低问题维度提升建模效率。本文基于最近的流匹配与权重空间图神经网络工作,设计了三种不同的权重空间流模型。每种模型采用不同方式刻画权重几何结构,系统探索了权重空间流的设计空间。实验表明,更忠实建模神经网络几何结构的模型更具有效性,能跨任务和架构泛化;我们的方法在参数量减少一个数量级的情况下仍具竞争力,并可通过规模扩展进一步提升性能。最后,我们列出未来研究方向。
原文摘要 · Abstract (English)
Deep generative models such as flow and diffusion models have proven to be effective in modeling high-dimensional and complex data types such as videos or proteins, and this has motivated their use in different data modalities, such as neural network weights. A generative model of neural network weights would be useful for a diverse set of applications, such as Bayesian deep learning, learned optimization, and transfer learning. However, the existing work on weight-space generative models often ignores the symmetries of neural network weights, or only takes into account a subset of them. Modeling those symmetries, such as permutation symmetries between subsequent layers in an MLP, the filters in a convolutional network, or scaling symmetries arising with the use of non-linear activations, holds the potential to make weight-space generative modeling more efficient by effectively reducing the dimensionality of the problem. In this light, we aim to design generative models in weight-space that more comprehensively respect the symmetries of neural network weights. We build on recent work on generative modeling with flow matching, and weight-space graph neural networks to design three different weight-space flows. Each of our flows takes a different approach to modeling the geometry of neural network weights, and thus allows us to explore the design space of weight-space flows in a principled way. Our results confirm that modeling the geometry of neural networks more faithfully leads to more effective flow models that can generalize to different tasks and architectures, and we show that while our flows obtain competitive performance with orders of magnitude fewer parameters than previous work, they can be further improved by scaling them up. We conclude by listing potential directions for future work on weight-space generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。