突破基准测试准确率瓶颈,多维度评估智能体性能
Life After Benchmark Saturation: A Case Study of CORE-Bench

- 提出六维评估框架,超越单一准确率
- 在核心基准上发现隐蔽的思维捷径问题
- 验证人类协作可实现约2倍效率提升
当基准测试准确率趋于饱和时,通常会被更难的新版本替代。我们指出这一做法过度关注准确率,忽视了智能体性能的六个关键维度:构念效度问题(如捷径)、分布外泛化能力、效率、可靠性、模型与支架的相对贡献,以及人机协作带来的提升。以计算可复现性任务的CORE-Bench Hard为案例,我们展示即使准确率饱和,沿这些维度评估仍能获得有意义的洞察。首先,我们识别出难以被低能力智能体察觉的构念效度威胁,并推出改进版基准CORE-Bench v1.1及分布外任务套件CORE-Bench OOD。其次,尽管准确率已达饱和,CORE-Bench v1.1仍可用于衡量效率、可靠性、模型与支架表现。最后,我们开展小规模随机实验,测量真实任务中的人机协作增益。结果显示,协作带来约两倍的提速——可能低估,因五分之一的人类独立复现实验未在时限内完成。整体贡献提供了一种比主流准确率导向范式更严谨的评估路径。
原文摘要 · Abstract (English)
When a benchmark's accuracy saturates, it is often retired and replaced with a more challenging version. We show that this approach privileges accuracy and misses the opportunity to study six other key dimensions of agent performance: construct validity issues such as shortcuts, out-of-distribution generalizability, efficiency, reliability, the relative importance of the model versus the scaffold, and uplift from human-agent collaboration. We use CORE-Bench Hard, a benchmark for computational reproducibility of scientific code, as a case study to demonstrate that measuring agents along these dimensions yields meaningful insights into agent performance even after accuracy saturates. First, we surface threats to construct validity in CORE-Bench Hard that are difficult to anticipate with less capable agents. We introduce an improved benchmark, CORE-Bench v1.1, and an out-of-distribution task suite, CORE-Bench OOD. Second, we find that despite accuracy saturation, CORE-Bench v1.1 remains useful for measuring efficiency, reliability, model performance, and scaffold performance. Finally, we conduct a small-scale randomized experiment to measure uplift from human-agent collaboration on real-world computational reproducibility tasks. We find a statistically significant speedup by about a factor of two -- likely underestimated due to one-fifth of human-only reproductions reaching the time limit before completing -- and describe various other findings. Together, our contributions present a more rigorous alternative to the dominant accuracy-centric evaluation paradigm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。