arXiv:2510.05197cs.AIcs.LG2025-10被引 8

提出新方法,低成本精准预测大模型多次尝试后的表现

Efficient Prediction of Pass@k Scaling in Large Language Models

  • 用贝塔-二项分布改进数据不足时的预测精度
  • 动态分配采样预算,重点攻克难题,提升罕见事件预测能力
  • 适合模型厂商和监管机构评估大规模使用风险与能力

评估前沿AI系统的性能与风险是关键研究方向。已有研究表明,对模型进行重复采样可显著提升其能力(如解决复杂数学与编程问题)和潜在危害(如越狱攻击)。这引发一个重要问题:在仅能进行少量采样的情况下,如何准确预测模型在海量尝试下的表现?这对日均服务数亿用户的模型提供商及希望防范风险的监管机构至关重要。本文提出三项贡献:首先,发现现有拟合方法在数据有限时存在统计缺陷,影响预测准确性;其次,引入基于贝塔-二项分布的鲁棒估计框架,在数据稀缺下实现更精确预测;第三,设计动态采样策略,将更多采样资源分配给更难问题。三者结合,使稀有风险与能力的预测更可靠,且计算成本大幅降低。

原文摘要 · Abstract (English)

Assessing the capabilities and risks of frontier AI systems is a critical area of research, and recent work has shown that repeated sampling from models can dramatically increase both. For instance, repeated sampling has been shown to increase their capabilities, such as solving difficult math and coding problems, but it has also been shown to increase their potential for harm, such as being jailbroken. Such results raise a crucial question for both capability and safety forecasting: how can one accurately predict a model's behavior when scaled to a massive number of attempts, given a vastly smaller sampling budget? This question is directly relevant to model providers, who serve hundreds of millions of users daily, and to governmental regulators, who seek to prevent harms. To answer this questions, we make three contributions. First, we find that standard methods for fitting these laws suffer from statistical shortcomings that hinder predictive accuracy, especially in data-limited scenarios. Second, we remedy these shortcomings by introducing a robust estimation framework, which uses a beta-binomial distribution to generate more accurate predictions from limited data. Third, we propose a dynamic sampling strategy that allocates a greater budget to harder problems. Combined, these innovations enable more reliable prediction of rare risks and capabilities at a fraction of the computational cost.

大模型评估采样策略风险预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。