arXiv:2507.00885cs.CLcs.LG2025-07EMNLP被引 30

大多数下游任务无法可靠预测模型性能,实验设置微调就可能改变结果。

Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check

  • 通过元分析发现仅39%的下游任务存在可预测的缩放规律。
  • 实验设置的小变化会导致缩放行为完全改变,稳定性差。
  • 提醒研究者关注非线性趋势,避免盲目依赖线性外推。

下游缩放定律旨在从模型在小规模下的表现预测其在大规模下的任务性能。这种预测是否可行尚不明确:部分研究在对性能指标进行简单变换后发现了清晰的线性缩放趋势,而另一些研究则指出根本性挑战,如涌现现象和反向缩放。本文对现有下游缩放定律数据进行了元分析,发现可预测的缩放仅在少数情况下成立:39%的时间。此外,实验设置的细微变化可能导致缩放行为发生彻底改变。该分析强调了理解缩放定律适用条件的重要性。要准确建模预训练损失与任务性能之间的关系,必须接纳那些偏离线性趋势的案例。

原文摘要 · Abstract (English)

Downstream scaling laws aim to predict task performance at larger scales from the model's performance at smaller scales. Whether such prediction should be possible is unclear: some works discover clear linear scaling trends after simple transformations of the performance metric, whereas others point out fundamental challenges to downstream scaling laws, such as emergence and inverse scaling. In this work, we conduct a meta-analysis of existing data on downstream scaling laws, and we find that predictable scaling only occurs in a minority of cases: 39% of the time. Moreover, seemingly benign changes to the experimental setting can completely change the scaling behavior. Our analysis underscores the need to understand the conditions under which scaling laws succeed. To accurately model the relationship between pretraining loss and task performance, we must embrace the cases in which scaling behavior deviates from linear trends.

缩放定律模型评估实验可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。