arXiv:2512.05117cs.LGcs.AI2025-12被引 24

发现深度模型训练后会收敛到共享的低维权重子空间。

The Universal Weight Subspace Hypothesis

  • 通过分析千余模型,发现不同任务下权重矩阵存在共通的低维主方向。
  • 仅用少数主方向即可捕捉多数权重变化,且与初始化、任务无关。
  • 适合研究模型泛化、合并及高效训练的学者参考。

我们证明,跨多种任务训练的深度神经网络表现出显著相似的低维参数子空间。这是首个大规模实证研究,表明神经网络在不同初始化、任务和领域下均系统性收敛到共享的谱子空间。通过对超过1100个模型(包括500个Mistral-7B LoRAs、500个Vision Transformers和50个LLaMA-8B模型)进行逐模式谱分析,我们识别出仅用少数主方向即可捕获大部分方差的通用子空间。通过在多种架构上对权重矩阵进行谱分解,我们发现稀疏的联合子空间在不同任务和数据集间被一致利用。这些发现为深层网络内部信息组织提供了新视角,并引发关于无需大量数据与算力即可发现这些通用子空间的可能性的思考。该内在结构对模型可复用性、多任务学习、模型合并及高效训练推理算法具有重要意义,或有助于降低大规模神经模型的碳足迹。

原文摘要 · Abstract (English)

We show that deep neural networks trained across diverse tasks exhibit remarkably similar low-dimensional parametric subspaces. We provide the first large-scale empirical evidence that demonstrates that neural networks systematically converge to shared spectral subspaces regardless of initialization, task, or domain. Through mode-wise spectral analysis of over 1100 models - including 500 Mistral-7B LoRAs, 500 Vision Transformers, and 50 LLaMA-8B models - we identify universal subspaces capturing majority variance in just a few principal directions. By applying spectral decomposition techniques to the weight matrices of various architectures trained on a wide range of tasks and datasets, we identify sparse, joint subspaces that are consistently exploited, within shared architectures across diverse tasks and datasets. Our findings offer new insights into the intrinsic organization of information within deep networks and raise important questions about the possibility of discovering these universal subspaces without the need for extensive data and computational resources. Furthermore, this inherent structure has significant implications for model reusability, multi-task learning, model merging, and the development of training and inference-efficient algorithms, potentially reducing the carbon footprint of large-scale neural models.

权重子空间模型合并泛化能力高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。