arXiv:2608.22145cs.LGcs.NA2026-08

揭示Adam在复杂损失曲面中失效的根源并提出可量化指标。

Loss Landscape Features That Make Adam Stall: Definitions, Estimators, and the Preconditioned Hessian View

论文配图:Loss Landscape Features That Make Adam Stall: Definitions, Estimators, and the Preconditioned Hessian View
图 1 · 摘自论文原文
  • 通过预条件海森矩阵等指标分析Adam失效机制
  • 实测显示二阶方法比Adam提升120-134dB图像精度
  • 适用于研究优化器性能或高精度图像重建的读者

在隐式神经表示架构和解析基准上,我们发现经过充分调参的Adam(尤其是学习率,在0.05到10^-8范围内扫参)即使在病态损失曲面上仍可能达到极低损失,或在远高于二阶方法的平台处收敛。本报告定义了判断Adam能否缓解特定损失曲面病态性的度量指标:海森矩阵的条件数及其亚当预条件版本D^{-1/2}HD^{-1/2}(由亚当更新规则推导)、区分轴对齐与交叉耦合病态性的对角质量ρ、由随机兰奇斯二次型估计的负谱质量,以及沿曲率带划分的梯度能量占比,包括指示亚当停滞的平坦区域占比。通过一个2×2示例与图示说明,亚当的对角预条件可重标定轴对齐病态性,但对交叉耦合病态性无效。此外,以FINER图像拟合架构为例,完整呈现了损失曲面分析框架:架构描述、其导致亚当在鞍点停滞的原因、经调优基线测得的峰值信噪比结果,及块状二阶方法实现的120–134 dB性能,附带误差图,并说明此类高精度图像拟合的实际优势。

原文摘要 · Abstract (English)

Across implicit-neural-representation (INR) architectures and analytic benchmarks we observe that a thoroughly tuned Adam (especially its learning rate (lr), e.g. in a hyperparameter sweep from $lr = 0.05$ to $10^{-8}$) can potentially reach a very low loss even on ill-conditioned loss landscape or converge at a plateau far above the loss attained by second-order methods. This report defines the measured metrics that help determine if Adam can mitigate the ill-conditioning on a given loss landscape. We provide the indicators by which each outcome is determined, that are: the condition number of the Hessian and of the Adam-preconditioned Hessian $D^{-1/2}HD^{-1/2}$ (with the derivation from Adam's update rule), the diagonal mass $ρ$ that distinguishes axis-aligned from cross-coupled ill-conditioning, the negative spectral mass estimated by stochastic Lanczos quadrature, and the gradient energy fractions over curvature bands, including the flat fraction that indicates the Adam stall. A worked out $2\times 2$ example and an illustration show the reasons why a diagonal preconditioning by Adam can remove axis-aligned ill-conditioning by rescaling and why it cannot do the same if the ill-conditioning is cross coupled. In addition, we present a case study of FINER image fitting architecture that goes over the whole loss landscape analysis framework: the fitting architecture description, reasons due to which its landscape stalls Adam at saddles, the measured PSNR values through our tuned baselines to the $120$--$134$\,dB results of the blockwise second order methods, the error maps behind those numbers, and description of the benefits such image fitting accuracy gives in practice.

优化器分析损失曲面图像重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。