用费舍尔几何优化神经网络不确定性预测,提升稳定性与准确性。
Fisher8: Stabilizing Neural Heteroscedastic Regression via Output-Layer Fisher Geometry

- 基于输出层费舍尔几何修正梯度方向,更贴合损失曲面局部结构。
- 在多维回归与表征学习中,显著提升似然与误差权衡表现。
- 无需额外数据依赖超参,适合需要可靠不确定性的模型应用。
从噪声观测中联合预测均值和不确定性估计的神经网络训练常不稳定,引发多项独立的稳定化尝试。本文认为这些方法揭示了共性问题:梯度更新与损失曲面几何不匹配。为此,提出Fisher8,一种基于输出层费舍尔几何的梯度修正方法,通过重定向并缩放更新以匹配局部曲率,而非采用欧几里得几何。与以往稳定器不同,Fisher8仅需学习率作为超参数,且可近似计算连续预测分布间的KL信任半径。我们证明先前稳定器均收敛于该几何修正的重叠部分。在多维回归与表征学习任务中,Fisher8实现了更优的似然-误差权衡,预测出校准良好的不确定性,并学习到丰富的不确定性感知特征空间。
原文摘要 · Abstract (English)
Training neural networks to jointly predict mean and uncertainty estimates from noisy observations can be unstable, prompting a series of independent stabilization efforts. We argue that these interventions highlight a common underlying issue where gradient steps are poorly aligned with the geometry of the loss landscape. To better align updates with local curvature, we derive Fisher8, an output-layer gradient correction that reorients and rescales updates using Fisher geometry rather than Euclidean geometry. Unlike past stabilizers, Fisher8 introduces no data-dependent hyperparameters beyond learning rate and admits an approximate KL trust radius between successive predictive distributions. We show that prior stabilizers converge on overlapping components of this geometric correction. Across multidimensional regression and representation-learning tasks, Fisher8 obtains superior likelihood--error tradeoffs, predicts calibrated uncertainty estimates, and learns rich uncertainty-aware feature spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。