arXiv:2410.05107cs.LG2024-10中稿 · University of St被引 1

从神经网络权重中学习通用表征,实现模型性质的预测与生成。

Hyper-Representations: Learning from Populations of Neural Networks

  • 提出超表征方法,自监督学习神经网络权重的通用表征。
  • 可识别模型性能、训练状态和超参数等属性,支持定向采样生成。
  • 跨模型规模、架构和任务泛化,适用于微调与迁移学习。

本论文从神经网络最基础的组成部分——权重出发,探讨其蕴含的可学习结构。核心贡献是提出超表征(hyper-representations),一种自监督方法,用于学习神经网络权重的通用表征。实验表明,训练后的神经网络在权重空间中占据有意义的结构,这些结构可被学习并利用。超表征能揭示模型的性能、训练状态和超参数等属性,并支持在表征空间中定位特定性质区域,从而采样和生成具有目标特性的模型权重。该方法在微调和迁移学习中取得显著成功。此外,研究展示了超表征可泛化至不同模型规模、架构和任务,为构建神经网络基础模型提供了可能,实现跨模型与架构的知识聚合与复用。本研究推动了对神经网络权重结构的深层理解,有助于发展更可解释、高效且灵活的模型。

原文摘要 · Abstract (English)

This thesis addresses the challenge of understanding Neural Networks through the lens of their most fundamental component: the weights, which encapsulate the learned information and determine the model behavior. At the core of this thesis is a fundamental question: Can we learn general, task-agnostic representations from populations of Neural Network models? The key contribution of this thesis to answer that question are hyper-representations, a self-supervised method to learn representations of NN weights. Work in this thesis finds that trained NN models indeed occupy meaningful structures in the weight space, that can be learned and used. Through extensive experiments, this thesis demonstrates that hyper-representations uncover model properties, such as their performance, state of training, or hyperparameters. Moreover, the identification of regions with specific properties in hyper-representation space allows to sample and generate model weights with targeted properties. This thesis demonstrates applications for fine-tuning, and transfer learning to great success. Lastly, it presents methods that allow hyper-representations to generalize beyond model sizes, architectures, and tasks. The practical implications of that are profound, as it opens the door to foundation models of Neural Networks, which aggregate and instantiate their knowledge across models and architectures. Ultimately, this thesis contributes to the deeper understanding of Neural Networks by investigating structures in their weights which leads to more interpretable, efficient, and adaptable models. By laying the groundwork for representation learning of NN weights, this research demonstrates the potential to change the way Neural Networks are developed, analyzed, and used.

权重表征自监督模型生成迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。