揭示权重空间网络的表达能力,发现其等价性并提升性能34%。
On the Expressive Power of Permutation-Equivariant Weight-Space Networks
- 证明主流排列等变网络表达能力相同
- 在温和假设下实现权重与函数空间的通用逼近
- 指导改进模型,性能超越现有最优34%
权重空间学习研究直接作用于其他神经网络参数的架构。随着预训练模型日益普及,近期工作已证实权重空间网络在多种任务上的有效性。当前最先进方法依赖排列等变设计以提升泛化能力,但可能损害表达能力,亟需理论分析。不同于其他结构化领域,权重空间学习需处理同时作用于权重和函数空间的映射,使表达力分析尤为复杂。尽管已有少量工作提供部分表达力结果,全面表征仍缺。本文系统构建了权重空间网络表达力的理论框架:首先证明主流排列等变网络表达力等价;随后在输入权重满足温和自然假设下,建立权重空间与函数空间的普遍性,并刻画普遍性失效的极端情形。基于理论指导,对现有模型稍作修改,即实现比先前最优提升34%,验证了框架的实际价值。
原文摘要 · Abstract (English)
Weight-space learning studies neural architectures that operate directly on the parameters of other neural networks. Motivated by the growing availability of pretrained models, recent work has demonstrated the effectiveness of weight-space networks across a wide range of tasks. SOTA weight-space networks rely on permutation-equivariant designs to improve generalization. However, this may negatively affect expressive power, warranting theoretical investigation. Importantly, unlike other structured domains, weight-space learning targets maps operating on both weight and function spaces, making expressivity analysis particularly subtle. While a few prior works provide partial expressivity results, a comprehensive characterization is still missing. In this work, we address this gap by developing a systematic theory for expressivity of weight-space networks. We first prove that all prominent permutation-equivariant networks are equivalent in expressive power. We then establish universality in both weight- and function-space settings under mild, natural assumptions on the input weights, and characterize the edge-case regimes where universality no longer holds. Guided by our theoretical results, we show that slight modifications to existing weight-space models yield a 34% improvement over prior SOTA, demonstrating the practical relevance of our framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。