arXiv:2505.12096cs.LGcs.AI2025-05

初始权重的偏差影响模型可训练性,好初始化反而有偏。

When Bias Meets Trainability: Connecting Theories of Initialization

  • 将初始化偏差与梯度行为理论统一,揭示网络先天偏好
  • 发现高效学习依赖于对特定类别的初始偏向
  • 适合研究模型初始化与泛化性能关系的学者

深度神经网络(DNN)在初始化时的统计特性对其可训练性及架构固有偏差至关重要。经典均场(MF)理论表明,随机初始化网络的参数分布强烈影响梯度行为,决定其是否爆炸或消失。近期研究发现,未训练的DNN存在初始猜测偏差(IGB),即输入空间的大片区域被分配给单一类别。本文首次为大量DNN提供理论证明,将IGB与先前的MF理论相连接,表明高效学习与网络对特定类别的初始偏好紧密相关。这一联系得出反直觉结论:优化可训练性的初始化系统性地带有偏差,而非中立。

原文摘要 · Abstract (English)

The statistical properties of deep neural networks (DNNs) at initialization play an important role to comprehend their trainability and the intrinsic architectural biases they possess before data exposure Well established mean field (MF) theories have uncovered that the distribution of parameters of randomly initialized networks strongly influences the behavior of the gradients, dictating whether they explode or vanish. Recent work has showed that untrained DNNs also manifest an initial guessing bias (IGB), in which large regions of the input space are assigned to a single class. In this work, we provide a theoretical proof that links IGB to previous MF theories for a vast class of DNNs, showing that efficient learning is tightly connected to a network prejudice towards a specific class. This connection leads to a counterintuitive conclusion: the initialization that optimizes trainability is systematically biased rather than neutral.

初始化可训练性偏差分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。