用模型表现推断人类完成任务的时间,实现跨基准的高效评估。
BRIDGE: Predicting Human Task Completion Time From Model Performance
- 基于模型输出构建潜空间难度量表,与人类耗时对齐。
- 任务难度与人类耗时对数呈线性关系,可预测新任务耗时。
- 适合评估AI能力演进,尤其关注规模化模型性能趋势的人看。
评估AI系统的真实能力需将基准表现转化为人类可理解的任务难度度量。现有依赖人工标注完成时间的方法成本高、噪声大且难以扩展。本文提出BRIDGE,一种统一的心理测量框架,通过模型响应学习潜空间难度,并锚定至人类任务完成时间。采用双参数逻辑项反应理论模型,联合估计任务难度与模型能力。实验表明,潜空间任务难度与人类完成时间的对数呈线性关系,仅凭模型表现即可推断新基准的任务耗时。基于此对齐,我们以人类任务长度预测前沿模型能力,并独立复现METR的指数级增长规律:50%可解任务的临界点每约6个月翻倍。
原文摘要 · Abstract (English)
Evaluating the real-world capabilities of AI systems requires grounding benchmark performance in human-interpretable measures of task difficulty. Existing approaches that rely on direct human task completion time annotations are costly, noisy, and difficult to scale across benchmarks. In this work, we propose BRIDGE, a unified psychometric framework that learns a latent difficulty scale from model responses and anchors it to human task completion time. Using a two-parameter logistic Item Response Theory model, we jointly estimate latent task difficulty and model capability from model performance data across multiple benchmarks. We demonstrate that latent task difficulty varies linearly with the logarithm of human completion time, allowing human task completion time to be inferred for new benchmarks from model performance alone. Leveraging this alignment, we forecast frontier model capabilities in terms of human task length and independently reproduce METR's exponential scaling results, with the 50% solvable task horizon doubling approximately every 6 months.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。