arXiv:2410.16377stat.MLcs.AI2024-10被引 39

用记忆机制解释大模型多次推理的性能提升规律。

A Simple Model of Inference Scaling Laws

  • 基于记忆假设,构建推理次数与成功率的关系模型。
  • 发现推理失败率随尝试次数呈幂律下降,符合观测数据。
  • 适合研究提示成本与推理效率的科研人员参考。

神经网络的缩放定律因能预测参数、数据和算力增加时的模型性能而受到广泛关注。本文提出一个基于记忆的简单统计假设,研究推理过程中的缩放规律,特别是性能如何随多次推理尝试而提升。我们探讨了覆盖度(coverage)或pass@k指标,该指标衡量在重复尝试中成功概率,并为大语言模型在推理任务中观察到的推理缩放行为函数形式提供了动机。我们定义了一种“推理损失”,其随尝试次数增加呈现幂律衰减,并将其与提示成本关联。通过在简单生成模型上的实验验证,发现我们的预测与受控环境下的实证覆盖曲线一致。该简单框架为将推理缩放与其他已知缩放定律结合奠定了基础。

原文摘要 · Abstract (English)

Neural scaling laws have garnered significant interest due to their ability to predict model performance as a function of increasing parameters, data, and compute. In this work, we propose a simple statistical ansatz based on memorization to study scaling laws in the context of inference, specifically how performance improves with multiple inference attempts. We explore the coverage, or pass@k metric, which measures the chance of success over repeated attempts and provide a motivation for the observed functional form of the inference scaling behavior of the coverage in large language models (LLMs) on reasoning tasks. We then define an "inference loss", which exhibits a power law decay as the number of trials increases, and connect this result with prompting costs. We further test our construction by conducting experiments on a simple generative model, and find that our predictions are in agreement with the empirical coverage curves in a controlled setting. Our simple framework sets the ground for incorporating inference scaling with other known scaling laws.

推理缩放大模型幂律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。