arXiv:2501.13734cs.LG2025-01NeurIPS被引 13

提出深度学习超参数调优的理论复杂度分析框架,解决隐式目标函数波动难题。

Sample complexity of data-driven tuning of model hyperparameters in neural networks with structured parameter-dependent dual function

  • 基于数据驱动设定,用微分几何工具分析超参数变化对目标函数的影响
  • 证明在特定任务分布下,超参数调优的样本复杂度有理论上限
  • 适用于激活函数插值和图神经网络核参数调优,适合理论研究者参考

现代机器学习算法,尤其是基于深度学习的方法,通常需要精细的超参数调优以获得最佳性能。尽管贝叶斯优化、随机搜索等自动化方法受到广泛关注,但深度神经网络超参数调优的根本学习理论复杂度仍不清晰。本文针对这一空白,引入新的数据驱动设定,研究在一系列深度学习任务上平均表现最优的超参数调优复杂度。核心挑战在于目标函数随超参数变化剧烈且隐式依赖于模型参数的优化问题。为此,我们提出一种新方法,刻画固定问题实例下目标函数在超参数变化时的间断与振荡特性,分析依赖微分/代数几何及约束优化工具。该方法表明对应目标函数族的学习理论复杂度有界。我们进一步给出具体应用的样本复杂度边界:在调节神经激活函数插值系数和图神经网络核参数时,可实现高效调优。

原文摘要 · Abstract (English)

Modern machine learning algorithms, especially deep learning based techniques, typically involve careful hyperparameter tuning to achieve the best performance. Despite the surge of intense interest in practical techniques like Bayesian optimization and random search based approaches to automating this laborious and compute intensive task, the fundamental learning theoretic complexity of tuning hyperparameters for deep neural networks is poorly understood. Inspired by this glaring gap, we initiate the formal study of hyperparameter tuning complexity in deep learning through a recently introduced data driven setting. We assume that we have a series of deep learning tasks, and we have to tune hyperparameters to do well on average over the distribution of tasks. A major difficulty is that the utility function as a function of the hyperparameter is very volatile and furthermore, it is given implicitly by an optimization problem over the model parameters. To tackle this challenge, we introduce a new technique to characterize the discontinuities and oscillations of the utility function on any fixed problem instance as we vary the hyperparameter; our analysis relies on subtle concepts including tools from differential/algebraic geometry and constrained optimization. This can be used to show that the learning theoretic complexity of the corresponding family of utility functions is bounded. We instantiate our results and provide sample complexity bounds for concrete applications tuning a hyperparameter that interpolates neural activation functions and setting the kernel parameter in graph neural networks.

超参数调优理论分析深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。