arXiv:2607.09739cs.AIcs.CL2026-07被引 1

用语义嵌入选小样本提示,低成本高效逼近大模型评测结果。

Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks

论文配图:Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks
图 1 · 摘自论文原文
  • 基于语义嵌入的子模函数选择提示子集,无需模型评分
  • 在35个基准上,设施选址函数比12种基线更保真分数与排名
  • 适用于无评估和有少量评估场景,计算成本更低

我们研究大模型评测基准的核集(coreset)选择:从多个基准中选出少量提示,使其产生的模型得分和排名近似于完整基准集。在评价无监督的基准核集选择方法中,算法不依赖模型输出结果,以细粒度方式在多个基准上生成提示子集,而非选择整个基准子集。采用子模优化,设计并评估多种子模函数,包括基于确定性点过程(DPP)、子模互信息及设施选址(FL)函数。在包含35个异构基准、5类能力维度、18个前沿大模型和超过6.1万条提示的新大规模评测套件上,发现仅使用廉价语义嵌入的设施选址函数,在不同核集预算下,优于12种基于评分或多样性的基线。此外,该目标不限于无监督场景:在仅有少量完整基准需选择且可获取大量模型评分时,该方法在MMLU和MTEB排行榜上达到或超越现有最优,同时计算成本显著降低。结果表明,子模性是基准压缩的强大可靠工具。

原文摘要 · Abstract (English)

We study LLM benchmark coreset selection: selecting a small subset of prompts over multiple benchmarks whose induced model scores and rankings approximate those obtained from the full benchmark suite. In evaluation-unsupervised benchmark coreset selection (our approach), the selection algorithm uses no model evaluation outcomes, and operates on a fine granularity by producing subsets of prompts over multiple benchmarks rather than producing a sub-collection of entire benchmarks. We use submodular subset selection, and we develop and evaluate many different submodular functions for this purpose, including determinantal point process (DPP) based approaches, submodular mutual information functions, and facility location-based functions. On a new large-scale suite of 35 heterogeneous benchmarks spanning five different capability categories, 18 frontier LLMs, and over 61K prompts, we find that the facility location (FL) function operating exclusively on inexpensive semantic prompt embeddings preserves LLM scores better than twelve separate score-based and diversity-based baselines, across a range of coreset budgets. Moreover, we show our proposed objective is not limited to the evaluation-unsupervised regime: in the setting where only a handful of whole benchmarks must be selected and a large amount of model scores are available, the same objective matches or outperforms state-of-the-art baselines on the MMLU and MTEB leaderboards, while being substantially cheaper to compute. Together, our results suggest that submodularity, in general, is a strong and reliable tool for benchmark compression.

大模型评测提示选择子模优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。