arXiv:2412.06540cs.LGcs.AI2024-12NeurIPS被引 27

用隐含技能建模大模型性能,跨家族精准预测多基准表现。

Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families

  • 基于低维隐含技能构建新尺度定律,融合多基准相关性。
  • 在12个主流基准上预测准确率超传统方法,误差更低。
  • 适合研究模型可扩展性与计算资源优化的学者使用。

大语言模型的尺度定律通常依据参数规模和训练数据预测性能,但不同模型家族的训练配置与数据处理差异导致性能波动,单一尺度定律难以通用。训练每族专属尺度定律又需反复训练多种规模模型。本文提出技能尺度定律(SSLaws,发音为Sloth),利用公开基准数据,假设模型性能由推理、指令遵循等低维隐含技能驱动,这些技能受模型规模和训练令牌数影响,但不同家族效率各异。该方法通过挖掘基准间的相关性,实现更精确且可解释的性能预测,无需为每族训练多个模型。我们提供了参数识别的理论结果,并在来自Open LLM Leaderboard v1/v2的12个主流基准上进行实证评估,结果表明Sloth能高精度预测大模型性能,揭示复杂下游任务、测试时计算增加及技能计算最优扩展的规律。

原文摘要 · Abstract (English)

Scaling laws for large language models (LLMs) predict model performance based on parameters like size and training data. However, differences in training configurations and data processing across model families lead to significant variations in benchmark performance, making it difficult for a single scaling law to generalize across all LLMs. On the other hand, training family-specific scaling laws requires training models of varying sizes for every family. In this work, we propose Skills Scaling Laws (SSLaws, pronounced as Sloth), a novel scaling law that leverages publicly available benchmark data and assumes LLM performance is driven by low-dimensional latent skills, such as reasoning and instruction following. These latent skills are influenced by computational resources like model size and training tokens, but with varying efficiencies across model families. Sloth exploits correlations across benchmarks to provide more accurate and interpretable predictions while alleviating the need to train multiple LLMs per family. We present both theoretical results on parameter identification and empirical evaluations on 12 prominent benchmarks, from Open LLM Leaderboard v1/v2, demonstrating that Sloth predicts LLM performance accurately and offers insights into scaling behaviors for complex downstream tasks, increased test-time compute, and compute-optimal scaling of skills.

大模型尺度定律性能预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。