arXiv:2605.10405cs.LG2026-05被引 2

用低秩分解预测模型表现,高效准确选出最优大模型。

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization

  • 结合多臂赌博机与低秩预测,动态选择评估对象
  • 通过双重稳健估计降低预测偏差,保证结果可信
  • 实测减少大量计算开销,适合资源有限的模型筛选

在固定基准上选择最佳大语言模型通常成本高昂,因需对每个模型在每条样本上进行完整评估。多臂赌博机(MAB)算法可通过顺序选择待评估的模型-样本对,减少无效评估。进一步地,可利用部分观测到的模型-样本评分矩阵进行低秩分解预测得分,以节省评估次数。然而,此类预测非真实值,可能产生偏差,导致错误识别最优模型。本文提出一种原则性框架,将廉价预测与MAB结合,同时保持统计有效性。具体而言,我们推导出每种模型性能的双重稳健估计量,利用低秩预测降低方差。该方法可在自适应选模、无放回采样条件下构建有效的有限样本置信区间。实证结果表明,该方法显著减少所需评估次数,在真实基准上实现可观的计算与成本节约,且能准确识别最佳模型。

原文摘要 · Abstract (English)

Selecting the best large language model (LLM) for a fixed benchmark is often expensive, since exhaustive evaluation requires running every model on every example. Multi-armed bandit (MAB) algorithms can reduce the number of LLM calls by sequentially selecting the next model-example pair to evaluate, thereby avoiding wasted evaluations on clearly underperforming models. Further savings can be achieved by predicting model scores from the partially observed model-example score matrix using low-rank factorization. However, such predictions are not ground truth: they can be biased and may therefore lead to incorrect identification of the best model. In this work, we propose a principled framework that combines MAB with cheap predicted scores without compromising statistical validity. Specifically, we derive doubly robust estimators of each model's performance that use the low-rank predictions to reduce variance. This enables the construction of valid finite-sample confidence intervals in our setting, where models are selected adaptively and examples are sampled without replacement. Empirical results on real-world benchmarks show that our approach reduces the number of required evaluations, yielding meaningful savings in compute and cost while accurately identifying the best-performing model.

大模型评估低秩分解多臂赌博机高效筛选

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。