提出信号与噪声框架,提升小规模模型评估的可靠性。
Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation
- 用信号与噪声量化评估基准质量,优化决策可信度。
- 改进指标或过滤噪声任务可降低误差,提升预测准确性。
- 适合构建/选择评估基准的研究者参考使用。
大型语言模型研发成本高昂,常依赖小规模实验在多任务评估套件上做决策。本文分析了使评估基准更可靠的特性,并提出改进方法。引入两个关键指标:信号(区分优劣模型的能力)与噪声(对训练步随机波动的敏感度)。实验表明,信号-噪声比更高的基准在小规模决策中更可靠,低噪声基准的缩放定律预测误差更低。为此提出三项干预措施:采用更具信号与噪声优势的评估指标(如困惑度替代准确率)、过滤噪声子任务以提升整体信噪比、对模型中间检查点输出取平均以减少噪声。研究基于30个基准和6000万至320亿参数的375个开源模型,生成包含90万次评估结果的新公开数据集,共2亿实例。
原文摘要 · Abstract (English)
Developing large language models is expensive and involves making decisions with small experiments, typically by evaluating on large, multi-task evaluation suites. In this work, we analyze specific properties which make a benchmark more reliable for such decisions, and interventions to design higher-quality evaluation benchmarks. We introduce two key metrics that show differences in current benchmarks: signal, a benchmark's ability to separate better models from worse models, and noise, a benchmark's sensitivity to random variability between training steps. We demonstrate that benchmarks with a better signal-to-noise ratio are more reliable when making decisions at small scale, and those with less noise have lower scaling law prediction error. These results suggest that improving signal or noise will lead to more useful benchmarks, so we introduce three interventions designed to directly affect signal or noise. For example, we propose that switching to a metric that has better signal and noise (e.g., perplexity rather than accuracy) leads to better reliability and improved scaling law error. We also find that filtering noisy subtasks, to improve an aggregate signal-to-noise ratio, leads to more reliable multi-task evaluations. We also find that averaging the output of a model's intermediate checkpoints to reduce noise leads to consistent improvements. We conclude by recommending that those creating new benchmarks, or selecting which existing benchmarks to use, aim for high signal and low noise. We use 30 benchmarks for these experiments, and 375 open-weight language models from 60M to 32B parameters, resulting in a new, publicly available dataset of 900K evaluation benchmark results, totaling 200M instances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。