arXiv:2510.02721cs.LGcs.CL2025-10中稿 · COLM被引 3

发现超参最优附近损失曲面简单,可高效评估模型极限性能。

Hyperparameter Loss Surfaces Are Simple Near their Optima

  • 通过随机搜索揭示最优区的低维结构特征。
  • 随机搜索最佳结果服从新分布,参数对应曲面本质特征。
  • 可估算最优性能置信区间,判断有效超参数量。

超参数显著影响模型能力;然而现代模型规模过大,难以进行广泛搜索。研究者通常基于对超参数的理解设计跨尺度有效训练方案。尽管重要,但现有工具极少用于理解超参数损失曲面。我们发现其新结构,并提出新理论以构建相应工具。损失曲面整体复杂,但在接近最优时,结构趋于简单,表现为少数基本特征,如有效维度和最佳可能损失。为揭示此渐近区域,我们开发一种基于随机搜索的新技术。在此区域内,随机搜索获得的最佳得分呈现新分布,其参数恰好对应渐近区域的曲面特征。由此推导出随机搜索的渐近规律,可解释并外推其收敛行为。这些新工具支持新分析,如最佳性能置信区间估计或有效超参数量判定。相关工具已开源于 https://github.com/nicholaslourie/opda。

原文摘要 · Abstract (English)

Hyperparameters greatly impact models' capabilities; however, modern models are too large for extensive search. Instead, researchers design recipes that train well across scales based on their understanding of the hyperparameters. Despite this importance, few tools exist for understanding the hyperparameter loss surface. We discover novel structure in it and propose a new theory yielding such tools. The loss surface is complex, but as you approach the optimum simple structure emerges. It becomes characterized by a few basic features, like its effective dimension and the best possible loss. To uncover this asymptotic regime, we develop a novel technique based on random search. Within this regime, the best scores from random search take on a new distribution we discover. Its parameters are exactly the features defining the loss surface in the asymptotic regime. From these features, we derive a new asymptotic law for random search that can explain and extrapolate its convergence. These new tools enable new analyses, such as confidence intervals for the best possible performance or determining the effective number of hyperparameters. We make these tools available at https://github.com/nicholaslourie/opda .

超参优化损失曲面随机搜索渐近分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。