用大模型提升神经网络权重生成效率与泛化能力
Foundation Models Secretly Understand Neural Network Weights: Enhancing Hypernetwork Architectures with Foundation Models
- 用Transformer架构的大模型生成其他网络的权重参数
- 在多种任务中提升性能、泛化性和数据利用率
- 适合研究神经网络压缩与自适应生成的学者
大型预训练模型(即基础模型)在适配多种下游任务时表现出色,常优于专用模型。超网络是一种生成另一神经网络部分或全部参数的神经网络,已成为条件化和泛化隐式神经表示(INRs)的重要方法,用于以神经网络形式表示音频、3D形状等信号或对象。然而,尽管将基础模型引入超网络具有潜力,这一方向尚未被系统研究,可能是因为权重生成任务与其他视觉任务差异较大。为此,本文(1)展示了基于Transformer架构的基础模型如何改进超网络;(2)通过可泛化INR任务的实证分析,证明利用基础模型能显著提升多种算法和模态下的性能、泛化性与数据效率。此外,我们还深入探讨了基础模型驱动超网络的设计空间,包括基础模型选择、算法配置以及扩大基础模型的影响。
原文摘要 · Abstract (English)
Large pre-trained models, or foundation models, have shown impressive performance when adapted to a variety of downstream tasks, often out-performing specialized models. Hypernetworks, neural networks that generate some or all of the parameters of another neural network, have become an increasingly important technique for conditioning and generalizing implicit neural representations (INRs), which represent signals or objects such as audio or 3D shapes using a neural network. However, despite the potential benefits of incorporating foundation models in hypernetwork methods, this research direction has not been investigated, likely due to the dissimilarity of the weight generation task with other visual tasks. To address this gap, we (1) show how foundation models can improve hypernetworks with Transformer-based architectures, (2) provide an empirical analysis of the benefits of foundation models for hypernetworks through the lens of the generalizable INR task, showing that leveraging foundation models improves performance, generalizability, and data efficiency across a variety of algorithms and modalities. We also provide further analysis in examining the design space of foundation model-based hypernetworks, including examining the choice of foundation models, algorithms, and the effect of scaling foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。