多数有效学习规则本质都是自然梯度下降。
Is All Learning (Natural) Gradient Descent?
- 将学习规则重写为带度量的自然梯度形式。
- 找到可最小化条件数的最优度量矩阵。
- 适用于连续、离散、随机等多种学习场景。
本文表明,一类广泛有效的学习规则——在给定时间窗口内提升标量性能指标的规则——均可重写为相对于某个特定损失函数和度量的自然梯度下降。具体而言,这类学习规则的参数更新可表示为一个对称正定矩阵(即度量)与损失函数负梯度的乘积。我们还证明了这些度量具有规范形式,并识别出若干最优度量,包括能实现最小可能条件数的度量。主要结论的证明仅依赖于初等线性代数和微积分,适用于连续时间、离散时间、随机及高阶学习规则,以及显式依赖时间的损失函数。
原文摘要 · Abstract (English)
This paper shows that a wide class of effective learning rules -- those that improve a scalar performance measure over a given time window -- can be rewritten as natural gradient descent with respect to a suitably defined loss function and metric. Specifically, we show that parameter updates within this class of learning rules can be expressed as the product of a symmetric positive definite matrix (i.e., a metric) and the negative gradient of a loss function. We also demonstrate that these metrics have a canonical form and identify several optimal ones, including the metric that achieves the minimum possible condition number. The proofs of the main results are straightforward, relying only on elementary linear algebra and calculus, and are applicable to continuous-time, discrete-time, stochastic, and higher-order learning rules, as well as loss functions that explicitly depend on time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。