把模型权重当数据,用生成方式批量造好模型
Position: Weight Space Should Be a First-Class Generative AI Modality
- 把神经网络权重看作可生成的数据,直接在权重空间合成新模型
- 生成的权重能达微调效果,但成本降几个数量级
- 适合想快速迭代模型或构建自动化AI系统的研究者
神经网络检查点已悄然成为大规模数据资源:数百万训练好的权重向量存在,各自编码了特定任务、领域和架构的知识。本文主张应将模型检查点视为第一类数据模态,并将权重空间的生成建模标准化为机器学习核心范式。近期进展表明,神经网络权重可按需合成,性能常媲美微调,且适应成本降低数个数量级。我们认为这些结果反映了深层结构事实:高性能模型位于由对称性、平坦性、模块性和共享子空间塑造的低维、高度结构化的权重空间区域。基于此观点,我们梳理现有方法为五阶段流程,综述已有实际应用,并澄清当前局限:适配器级与条件生成发展迅速,而无约束的前沿规模检查点合成仍待突破。目标是推动社区从‘为每项任务优化模型’转向‘从学习到的权重分布中采样模型’,加速迈向AI系统能持续自我改进甚至创造新AI系统的时代。
原文摘要 · Abstract (English)
Neural network checkpoints have quietly become a large-scale data resource: millions of trained weight vectors now exist, each encoding task-, domain-, and architecture-specific knowledge. This position paper argues that model checkpoints should be treated as a first-class data modality, and that generative modeling in weight space should be standardized as a core machine learning primitive. Recent advances demonstrate that neural weights can be synthesized on demand, often matching fine-tuning performance while reducing adaptation cost by orders of magnitude. We contend that these results reflect an underlying structural fact: high-performing models occupy low-dimensional, highly structured regions of weight space shaped by symmetry, flatness, modularity, and shared subspaces. Building on this view, we organize existing methods into a five-stage pipeline, survey applications where the approach is already practical, and clarify current limits: adapter-scale and conditional generation are advancing rapidly, while unrestricted frontier-scale checkpoint synthesis remains open. Our goal is to shift the community's default mindset from optimizing models per task to sampling models from learned weight distributions, accelerating toward an era in which AI systems routinely improve or create other AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。