发现深度回归模型的性能随数据和模型规模呈幂律提升,效果显著。
Neural Scaling Laws for Deep Regression
- 通过多种网络结构验证了损失与数据量、模型容量的幂律关系。
- 缩放指数在1到2之间,表明加大数据能显著提升模型表现。
- 适合关注模型扩展规律与资源分配的研究者参考。
神经缩放定律——即泛化误差与深度学习模型特性之间的幂律关系——是高效构建可靠模型的重要工具。尽管大语言模型的成功凸显了其重要性,但该定律在深度回归模型中的应用仍鲜有探索。本文基于扭曲范德华磁体的参数估计任务,实证研究了深度回归中的神经缩放定律。结果表明,在广泛参数范围内,损失与训练数据规模及模型容量均呈现幂律关系,涵盖全连接网络、残差网络与视觉变换器等多种架构。控制这些关系的缩放指数介于1至2之间,具体值取决于被回归参数和模型细节。一致的缩放行为及其较大的指数表明,增加数据规模可显著改善深度回归模型的性能。
原文摘要 · Abstract (English)
Neural scaling laws--power-law relationships between generalization errors and characteristics of deep learning models--are vital tools for developing reliable models while managing limited resources. Although the success of large language models highlights the importance of these laws, their application to deep regression models remains largely unexplored. Here, we empirically investigate neural scaling laws in deep regression using a parameter estimation model for twisted van der Waals magnets. We observe power-law relationships between the loss and both training dataset size and model capacity across a wide range of values, employing various architectures--including fully connected networks, residual networks, and vision transformers. Furthermore, the scaling exponents governing these relationships range from 1 to 2, with specific values depending on the regressed parameters and model details. The consistent scaling behaviors and their large scaling exponents suggest that the performance of deep regression models can improve substantially with increasing data size.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。