arXiv:2509.20645cs.CLcs.AI2025-09被引 2

让模型提前预测自身表现,减少试错成本。

Anticipatory Evaluation of Language Models

  • 仅根据任务描述和配置,预测模型性能。
  • 最高预测误差为9.9,高置信度下仍具实用价值。
  • 适合研究评估效率与资源优化的学者。

大语言模型的发展正面临评估瓶颈:必须构建基准并运行模型后才能迭代。我们探究是否可在实验前就预测评估结果。具体研究文本仅性能预测,即模型仅凭任务描述和实验配置即可估计性能,无需接触数据样本。为支持系统性研究,我们构建了PRECOG数据集,包含跨越多种任务、领域和指标的描述-性能配对。从arXiv中抓取任务与配置描述,共获得2,290个实例,覆盖1,519篇论文,并使用模型知识截止日期后的论文构建测试集。实验表明该任务具有挑战性但可行:推理模型在高置信度阈值下达到最低均方绝对误差9.9。整体上,我们的数据集与分析为开放式的前瞻性评估提供了初步基础,有助于难度预估与更智能的资源分配。

原文摘要 · Abstract (English)

Progress in large language models is increasingly constrained by an evaluation bottleneck: benchmarks must be built and models run before iteration can begin. We investigate whether evaluation outcomes can be forecast before any experiments are conducted. Specifically, we study text-only performance prediction, where models estimate performance from task descriptions and experimental configurations alone, without access to dataset instances. To support systematic study, we curate PRECOG, a corpus of description-performance pairs spanning diverse tasks, domains, and metrics. We scrape task and configuration descriptions from arXiv, yielding 2,290 instances covering 1,519 papers, and construct a test split using papers published after the evaluated models' knowledge cutoff. Experiments show the task is challenging but feasible: reasoning models achieve a non-trivial forecasting skill reaching mean absolute error as low as 9.9 at high-confidence thresholds. Overall, our corpus and analyses offer an initial step toward open-ended anticipatory evaluation, supporting difficulty estimation and smarter resource allocation.

模型评估预测效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。