arXiv:2603.02622cs.LGstat.ML2026-03

揭示深度线性判别分析的隐式正则化机制

Implicit Bias in Deep Linear Discriminant Analysis

  • 通过梯度流分析发现权重更新呈乘法形式
  • 平衡初始化下自动保持(2/L)准范数守恒
  • 为度量学习优化几何提供理论支撑

尽管标准损失函数的隐式偏差已被研究,但判别性度量学习目标所诱导的优化几何仍缺乏系统探索。据我们所知,本文首次对深度线性判别分析(Deep LDA)——一种尺度不变的目标函数,旨在最小化类内方差并最大化类间距离——进行了理论分析。通过对L层对角线性网络的梯度流进行分析,证明在平衡初始化条件下,网络架构将标准加法型梯度更新转化为乘法型权重更新,从而实现(2/L)准范数的自动守恒。

原文摘要 · Abstract (English)

While the Implicit Bias(or Implicit Regularization) of standard loss functions has been studied, the optimization geometry induced by discriminative metric-learning objectives remains largely unexplored.To the best of our knowledge, this paper presents an initial theoretical analysis of the implicit regularization induced by the Deep LDA,a scale invariant objective designed to minimize intraclass variance and maximize interclass distance. By analyzing the gradient flow of the loss on a L-layer diagonal linear network, we prove that under balanced initialization, the network architecture transforms standard additive gradient updates into multiplicative weight updates, which demonstrates an automatic conservation of the (2/L) quasi-norm.

深度学习隐式正则化度量学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。