用图微分方程建模网络结构,更准预测训练曲线。
Architecture-Aware Learning Curve Extrapolation via Graph Ordinary Differential Equation
- 引入图微分方程融合网络结构信息建模训练过程
- 在MLP和CNN上均超越现有最先进方法
- 适合需要快速评估架构性能的AutoML场景
学习曲线外推可从早期训练阶段预测神经网络性能,广泛应用于加速自动化机器学习,助力超参数调优与神经网络架构搜索。然而,现有方法通常孤立建模学习曲线演变,忽视了神经网络架构对损失曲面和学习轨迹的影响。本文探索将架构信息融入学习曲线建模的可能性及有效整合方式。受优化过程动力系统视角启发,提出一种新颖的架构感知神经微分方程模型,实现学习曲线的连续预测。实验表明,该模型能有效捕捉波动学习曲线的整体趋势,并通过变分参数量化不确定性。其在MLP与基于CNN的学习曲线外推任务中均优于当前最先进方法及纯时间序列建模方法。此外,我们验证了该方法在神经架构搜索中用于训练配置排序等场景的适用性。
原文摘要 · Abstract (English)
Learning curve extrapolation predicts neural network performance from early training epochs and has been applied to accelerate AutoML, facilitating hyperparameter tuning and neural architecture search. However, existing methods typically model the evolution of learning curves in isolation, neglecting the impact of neural network (NN) architectures, which influence the loss landscape and learning trajectories. In this work, we explore whether incorporating neural network architecture improves learning curve modeling and how to effectively integrate this architectural information. Motivated by the dynamical system view of optimization, we propose a novel architecture-aware neural differential equation model to forecast learning curves continuously. We empirically demonstrate its ability to capture the general trend of fluctuating learning curves while quantifying uncertainty through variational parameters. Our model outperforms current state-of-the-art learning curve extrapolation methods and pure time-series modeling approaches for both MLP and CNN-based learning curves. Additionally, we explore the applicability of our method in Neural Architecture Search scenarios, such as training configuration ranking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。