统一模型预测深度学习性能,多维度协同变化下更准。
Unified Neural Scaling Laws

- 提出统一神经网络缩放定律,同时建模参数、数据量等多因素影响。
- 在视觉、语言、数学、强化学习任务中,外推精度显著优于现有方法。
- 适合研究模型规模与计算资源关系的科研人员参考。
我们提出一种函数形式(称为统一神经网络缩放定律,UNSL),能准确建模并外推深度神经网络在多个维度同时变化时的缩放行为——即当模型参数数量、训练数据集大小、训练步数、推理步数、计算量及各类超参数同时变化时,目标任务的评估指标如何演变。该模型适用于多种架构和广泛的上游与下游任务,涵盖大规模视觉、语言、数学与强化学习。相较于其他神经网络缩放函数形式,UNSL在该数据集上的缩放行为外推结果更加精确。
原文摘要 · Abstract (English)
We present a functional form (that we refer to as a Unified Neural Scaling Law (UNSL)) that accurately models and extrapolates the scaling behaviors of deep neural networks as multiple dimensions all vary simultaneously (i.e. how the evaluation metric of interest varies as one simultaneously varies the number of model parameters, training dataset size, number of training steps, number of inference steps, amount of compute, and various hyperparameters) for various architectures and for each of various tasks within a varied set of upstream and downstream tasks. This set includes large-scale vision, language, math, and reinforcement learning. When compared to other functional forms for neural scaling, this functional form yields extrapolations of scaling behavior that are considerably more accurate on this set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。