arXiv:2410.09973stat.MLcs.LG2024-10被引 2

高维下梯度跨度算法表现稳定,解释了大模型训练为何每次结果相似。

Gradient Span Algorithms Make Predictable Progress in High Dimension

  • 提出梯度跨度算法在高维随机函数中趋于确定性行为
  • 不同初始化下训练成本曲线几乎相同,验证可预测性
  • 适合关注自动化机器学习和减少重复训练的研究者

我们证明,所有‘梯度跨度算法’在维度趋于无穷时,对缩放的高斯随机函数具有渐近确定性行为。这是对随机二次函数和自旋玻璃类似结果的功能推广。该结论解释了大规模机器学习模型在复杂非凸景观中,尽管随机初始化,多次训练运行仍产生近似相同的代价曲线这一反直觉现象。这种‘可预测进展’已被AutoML社区利用:由于单次运行的优化轨迹已具代表性,无需用相同超参数重复多次尝试。

原文摘要 · Abstract (English)

We prove that all 'gradient span algorithms' have asymptotically deterministic behavior on scaled Gaussian random functions as the dimension tends to infinity. This is a functional generalization of similar results for random quadratic functions and spin glasses. They explain the counterintuitive phenomenon that different training runs of many large machine learning models result in approximately equal cost curves despite random initialization on a complicated non-convex landscape. This 'predictable progress' phenomenon is exploited by the AutoML community: Since the optimization progress of a single run is already representative, multiple retries with the same hyperparameters are not necessary.

优化理论机器学习高维优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。