arXiv:2504.10231cs.LG2025-04中稿 · ICLR被引 4

首个视觉Transformer模型动物园,助力模型分析与生成研究。

A Model Zoo of Vision Transformers

  • 构建包含250个独特ViT模型的系统化数据集
  • 覆盖多样训练策略,验证了权重空间与行为多样性
  • 适合从事模型分析、生成与迁移学习的研究者使用

大规模结构化神经网络集合(即‘模型动物园’)已推动模型分析、基于模型权重的表示学习及参数生成建模等下游任务的发展。然而,现有模型动物园规模有限,且忽视了当前最成功的神经网络架构之一——Transformer。本文首次提出视觉Transformer(ViT)模型动物园,设计新生成蓝图以涵盖预训练与微调全过程,发布250个精心生成的独特模型。这些模型覆盖广泛的生成因子,其多样性通过权重空间与行为指标全面验证。为展示数据集价值,我们结合探索性实验与文献案例,提出多种应用方向。该模型动物园使基于模型种群的方法从小型模型扩展至前沿架构,代码与模型已开源至github.com/ModelZoos/ViTModelZoo。

原文摘要 · Abstract (English)

The availability of large, structured populations of neural networks - called 'model zoos' - has led to the development of a multitude of downstream tasks ranging from model analysis, to representation learning on model weights or generative modeling of neural network parameters. However, existing model zoos are limited in size and architecture and neglect the transformer, which is among the currently most successful neural network architectures. We address this gap by introducing the first model zoo of vision transformers (ViT). To better represent recent training approaches, we develop a new blueprint for model zoo generation that encompasses both pre-training and fine-tuning steps, and publish 250 unique models. They are carefully generated with a large span of generating factors, and their diversity is validated using a thorough choice of weight-space and behavioral metrics. To further motivate the utility of our proposed dataset, we suggest multiple possible applications grounded in both extensive exploratory experiments and a number of examples from the existing literature. By extending previous lines of similar work, our model zoo allows researchers to push their model population-based methods from the small model regime to state-of-the-art architectures. We make our model zoo available at github.com/ModelZoos/ViTModelZoo.

视觉Transformer模型动物园生成建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。