arXiv:2507.10613cs.LGcs.AI2025-07被引 2

发现大模型性能提升放缓主因是数据冗余与资源分配不当

Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs

  • 通过400多个模型实证,发现高数据密度导致性能收益递减
  • 提出新缩放定律,能更准确预测大模型在低效训练时的表现
  • 适合关注模型效率、数据质量优化的研究者和工程师

自然语言处理中的传统缩放定律认为,增大模型规模和训练数据可提升性能。然而,近期研究显示,在大语言模型中性能提升逐渐放缓,这一现象称为次缩放。本文通过超过400个模型的广泛实证分析,重新审视缩放定律,揭示数据密度过高和资源分配非最优是导致次缩放的关键因素。高数据密度引发冗余信息,造成边际收益递减;而最优资源分配对持续性能提升至关重要。我们提出一种次非最优缩放定律,能更好预测次缩放阶段的模型表现,强调数据质量与多样性的关键作用。

原文摘要 · Abstract (English)

Traditional scaling laws in natural language processing suggest that increasing model size and training data enhances performance. However, recent studies reveal deviations, particularly in large language models, where performance improvements decelerate, which is a phenomenon known as sub-scaling. This paper revisits these scaling laws by examining the impact of data quality and training strategies on model performance. Through extensive empirical analysis of over 400 models, we identify high data density and non-optimal resource allocation as key factors contributing to sub-scaling. High data density leads to diminishing returns due to redundant information, while optimal resource allocation is crucial for sustained performance improvements. We propose a sub-optimal scaling law that better predicts performance in sub-scaling regimes, highlighting the importance of data quality and diversity.

大模型数据密度缩放定律训练策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。