arXiv:2501.19149cs.LGcs.AI2025-01被引 3

深度残差网络倾向于最小化瓶颈秩,提升泛化能力。

On the inductive bias of infinite-depth ResNets and the bottleneck rank

  • 通过分析无限深线性残差网络,发现其权重具有低秩偏好。
  • 在合适超参数下,非线性残差网络也趋向于最小化瓶颈秩。
  • 适合关注模型泛化机制的研究者阅读。

我们计算了深度线性残差网络的最小范数权重,发现该架构的归纳偏置介于最小化核范数和最小化秩之间。这意味着,在适当超参数下,深度非线性残差网络具有最小化瓶颈秩的归纳偏置。

原文摘要 · Abstract (English)

We compute the minimum-norm weights of a deep linear ResNet, and find that the inductive bias of this architecture lies between minimizing nuclear norm and rank. This implies that, with appropriate hyperparameters, deep nonlinear ResNets have an inductive bias towards minimizing bottleneck rank.

残差网络归纳偏置瓶颈秩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。