arXiv:2606.04212cs.LGstat.ML2026-06

边缘稳定态会不均等地分配学习进度,影响不同数据组的训练效果。

Edge of Stability Selectively Shapes Learning Across the Data Distribution

论文配图:Edge of Stability Selectively Shapes Learning Across the Data Distribution
图 1 · 摘自论文原文
  • 通过干预触发边缘稳定态,发现其选择性增强部分数据组的学习
  • 仅当梯度方向与海森矩阵主特征向量对齐且梯度不衰减时才获益
  • 适用于研究优化机制或数据分布偏移问题的研究者

现有对边缘稳定态(EoS)的分析将其视为优化的全局属性。我们证明它也具有选择性:稳定性约束会重新分配训练分布中不同子集的学习进度,增强某些群体的进展,同时抑制其他群体。通过从相同训练状态进入或退出EoS区间的分支干预,我们因果地展示了这一权衡,并确定了群体受益的两个必要条件:第一,其平均梯度必须与海森矩阵的主特征向量对齐;我们通过控制扰动验证该机制——保持距离但随机化方向,破坏对齐后优势消失。第二,该群体必须在时间上维持非零梯度幅值。在交叉熵损失下,分类置信度高的群体出现梯度饱和,导致其学习解耦,优势转移至输出异常点,其梯度持续存在。这些结果表明,EoS不仅是稳定性边界,更是调控学习在数据分布中分配的机制。

原文摘要 · Abstract (English)

Existing analyses of the edge of stability (EoS) treat it as a global property of optimization. We show that it is also selective: the stability constraint redistributes learning across subsets of the training distribution, amplifying progress on some groups while suppressing progress on others. Using a branching intervention that enters or exits the EoS regime from the same training state, we causally demonstrate this trade-off and identify two necessary conditions for a group to benefit. First, its aggregate gradient must align with the top Hessian eigenvector. We isolate this mechanism with a controlled perturbation that preserves distance but randomizes direction, destroying alignment and eliminating the advantage. Second, the group must sustain non-vanishing gradient magnitude over time. Under cross-entropy loss, gradient saturation decouples confidently classified groups, shifting the advantage to output-outliers, whose gradients persist. Together, these results show that EoS functions not only as a stability boundary, but as a mechanism governing the allocation of learning across the data distribution.

优化机制边缘稳定态学习分配梯度特性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。