梳理神经网络缩放定律的适用边界与实用策略
How to Upscale Neural Networks with Scaling Law? A Survey and Practical Guidelines
- 分析50+研究,提炼模型规模与数据量的预测关系
- 发现稀疏模型、多模态等架构常偏离传统缩放规律
- 适合关注大模型设计与部署优化的研究者阅读
神经网络缩放定律通过揭示模型规模、数据量与计算资源间的可预测关系,革新了大规模AI模型的设计与优化。早期研究确立了性能与规模之间的幂律关系,推动了计算最优的缩放策略。然而,近期研究表明,该规律在不同架构、模态和部署场景中存在局限:稀疏模型、专家混合、检索增强学习及多模态模型常偏离传统缩放模式。此外,视觉、强化学习与微调等领域的缩放行为差异显著,凸显需采用更精细的方法。本文综述超过50项研究,剖析缩放定律的理论基础、实证发现与实践意义,探讨数据效率、推理缩放及架构约束等关键挑战,主张根据实际应用定制适应性缩放策略。尽管缩放定律具指导价值,但其泛化能力并非普适。
原文摘要 · Abstract (English)
Neural scaling laws have revolutionized the design and optimization of large-scale AI models by revealing predictable relationships between model size, dataset volume, and computational resources. Early research established power-law relationships in model performance, leading to compute-optimal scaling strategies. However, recent studies highlighted their limitations across architectures, modalities, and deployment contexts. Sparse models, mixture-of-experts, retrieval-augmented learning, and multimodal models often deviate from traditional scaling patterns. Moreover, scaling behaviors vary across domains such as vision, reinforcement learning, and fine-tuning, underscoring the need for more nuanced approaches. In this survey, we synthesize insights from over 50 studies, examining the theoretical foundations, empirical findings, and practical implications of scaling laws. We also explore key challenges, including data efficiency, inference scaling, and architecture-specific constraints, advocating for adaptive scaling strategies tailored to real-world applications. We suggest that while scaling laws provide a useful guide, they do not always generalize across all architectures and training strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。