研究损失值能否作为衡量模型复杂度的坐标,发现只有特定路径才适用。
Loss-Parameterized Fisher Width Along Learning Trajectories
- 通过解析分解和稳定性分析,揭示了损失与参数空间几何的关系。
- 在逻辑回归模型中,损失越低时模型越接近最优解,且具有最大宽度。
- 不同优化器表现差异大,说明损失参数化需分路径而非通用适用。
Fisher 宽度衡量探针集经局部 Fisher 几何变形后的高斯宽度。本文研究其沿学习轨迹的演化,并探讨训练损失是否可作为该量的有效坐标。首先推导出精确的迹-形状分解及固定紧凑探针的确定性稳定性界。在种群高斯教师逻辑回归模型中,教师对齐态在损失低于 $\log 2$ 的所有水平上均为极值:参数范数最小,且最大化 Fisher 迹与欧氏球 Fisher 宽度。随后证明种群梯度流渐近选择此分支,并给出对齐与正交坐标的明确收敛速率。当 $d\geq2$ 时,有 \[ \frac{w_F(B_2^d;θ(t))} {\sqrt{L(θ(t))}} \longrightarrow \frac{\sqrt6}π\mathbb E[χ_{d-1}]. \] 控制全 Fisher 实验支持匹配损失分支与种群预测。在非线性 MLP 中使用对角模型 Fisher 近似,GD 和 SGD 在匹配损失下保持接近,而 Adam 则显著偏离;所测试的固定探针时间形状高度一致。结果支持基于分支而非普遍性的损失参数化方式。
原文摘要 · Abstract (English)
Fisher width measures the Gaussian width of a probe set after deformation by the local Fisher geometry. We study its evolution along learning trajectories and ask when training loss can serve as an effective coordinate for this quantity. We first derive an exact trace--shape factorization and a deterministic stability bound for fixed compact probes. In a population Gaussian-teacher logistic model, the teacher-aligned state is extremal on every loss level below $\log 2$: it has minimal parameter norm and maximizes both Fisher trace and Euclidean-ball Fisher width. We then show that population gradient flow asymptotically selects this branch, with explicit rates for the aligned and orthogonal coordinates. This yields, for $d\geq2$, \[ \frac{w_F(B_2^d;θ(t))} {\sqrt{L(θ(t))}} \longrightarrow \frac{\sqrt6}π\mathbb E[χ_{d-1}]. \] Controlled full-Fisher experiments support the matched-loss branch and the population predictions. In a nonlinear MLP with a diagonal model-Fisher approximation, GD and SGD remain close at matched loss, whereas Adam follows a substantially displaced branch; the fixed probes tested retain highly similar temporal shapes. These results support a branchwise, rather than universal, loss parametrization of Fisher width.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。