arXiv:2510.24626cs.CL2025-10被引 5

发现大模型规模扩大时,不同群体表现差距会分化而非缩小。

Relative Scaling Laws for LLMs

  • 用相同算力预算训练255个模型,比较不同数据分布的表现变化。
  • 学术题库趋同,方言差异随人口规模变化,AI风险行为出现分裂增长。
  • 适合关注模型公平性与安全性的研究者和开发者参考。

缩放定律描述了语言模型在增加数据、参数和计算量时的性能提升规律。尽管广泛应用,但通常基于聚合测试集进行评估,平均了异质子群体的表现,掩盖了性能差异。我们提出相对缩放定律,追踪不同测试分布间性能差距随规模演变的过程,而非仅关注绝对误差。在标准预训练数据集上,使用10^18至10^20 FLOPs的匹配算力(IsoFLOP)预算,训练了255个解码器型Transformer模型。结果表明:MMLU中的学术领域趋于表现均等;区域英语方言的表现变化取决于人口规模;人工智能风险行为呈现分化——能力与影响力相关风险在预训练中上升,而对抗性风险未见增长。这说明,虽然缩放提升了整体性能,却并非万能的均衡器。为支持后续研究,我们公开所有模型检查点,使从业者可同时测量相对与传统缩放定律,以更优地应对鲁棒性挑战。

原文摘要 · Abstract (English)

Scaling laws describe how language models improve with additional data, parameters, and compute. While widely used, they are typically measured on aggregate test sets. Aggregate evaluations yield clean trends but average over heterogeneous subpopulations, obscuring performance disparities. We introduce relative scaling laws, which track how performance gaps between test distributions evolve with scale rather than focusing solely on absolute error. Using 255 decoder-only Transformers trained under matched-compute (IsoFLOP) budgets from $10^{18}$--$10^{20}$ FLOPs on standard pretraining datasets, we find diverse trajectories: academic domains on MMLU converge toward parity; regional English dialects shift depending on population size; and clusters of AI risk behaviours split, with capability- and influence-related risks increasing during pretraining while adversarial risks do not. These results show that although scaling improves overall performance, it is not a universal equalizer. To support further study, we release all model checkpoints from this work to enable practitioners to measure relative alongside traditional scaling laws, in order to better prioritize robustness challenges in light of the bitter lesson.

大模型缩放定律公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。