arXiv:2606.26836cs.AI2026-06被引 1

现有评测低估大模型真实能力,新方法揭示82%性能差距

The Capability Frontier: Benchmarks Miss 82% of Model Performance

论文配图:The Capability Frontier: Benchmarks Miss 82% of Model Performance
图 1 · 摘自论文原文
  • 构建能力前沿模型,通过多模型多生成选择优化性能
  • 修正后准确率提升82%,顶尖模型成本降低85%
  • 适合关注真实场景评估与多模型协作的研究者

现有评测通常仅报告单个模型在单次运行中的准确率,系统性低估了大模型在异构数据分布下的真实能力:(i) 不同模型对不同问题表现各异,(ii) 在预算内可生成多个版本并择优保留。为此,我们提出能力前沿(Capability Frontier),即在最优跨模型与多生成选择下(即通过预言机),在每个成本水平上所能达到的最佳性能的帕累托前沿。该方法纠正了两种相反偏差:单一模型评估导致的低估,以及对噪声样本取最大值带来的高估。我们在16个主流基准上评估了21个大模型,涵盖编程、推理、医学、事实性、指令遵循和代理任务,对比了能力前沿与各基准最优模型在相同成本下的表现。纠正单模型评估使错误率降低54%;进一步纠正单次运行偏差,实现82%的性能提升,且顶级准确率可在85%的成本下达成。通过受控概率模拟,我们还发现更高查询主题熵会带来近似单调的性能差距扩大。结果表明,集体大模型能力被严重低估,这对数据异构、多领域场景下的评估与部署具有重要启示。

原文摘要 · Abstract (English)

Existing benchmarks typically report accuracy for a single model on a single run. This systematically understates real-world LLM capabilities, particularly under heterogeneous data distributions: (i) different models get different questions correct according to their specializations, and (ii) given a budget, multiple generations can be sampled and selectively retained. To quantify this gap, we introduce the Capability Frontier: a Pareto frontier over a set of models that characterizes the best achievable performance at each cost level under optimal selection across models and generations (i.e., via an oracle). Our construction corrects for two opposing biases: underestimation from single-model evaluation and overestimation from taking maxima over noisy samples. We study 21 LLMs across 16 widely used benchmarks spanning coding, reasoning, medicine, factuality, instruction following, and agentic tasks, comparing Capability Frontier performance at matched cost to each benchmark's top-performing model. Correcting for single-model evaluation yields a 54% error rate reduction; additionally correcting for single runs yields an 82% improvement, with SOTA accuracy matched at 85% cost reduction. Complementing these empirical results, we use controlled probabilistic simulations to show that higher query topic entropy produces a near-monotonic increase in the performance gap between oracle routing and the best single model. Our findings suggest collective LLM capabilities are substantially underestimated, with implications for evaluation and deployment in data-heterogeneous, multi-domain settings.

大模型评测能力前沿多模型协作性能评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。