arXiv:2510.25039cs.SEcs.LG2025-10被引 2

用AI自动设计动态评测基准,让模型评估更精准高效。

Automating Benchmark Design

  • 引入大模型辅助环境设计,参数化基准模板并自动搜索最优配置。
  • 生成的基准难度与目标偏差仅5.3%~13.2%,比基线提升2-4倍。
  • 适用于需要持续评估智能体能力的研究者和开发者。

大型语言模型(LLM)及其代理的快速发展超出了当前评估能力的范围。现有的人工设计静态基准很快就会饱和,而动态基准虽能随模型演进,但创建和维护成本高。为此,我们提出BeTaL(基于大模型回路的基准调优)框架,利用环境设计原则自动化动态基准设计。该方法通过参数化基准模板中的关键设计选项,借助大模型在参数空间中推理,以低成本实现目标属性(如难度、真实性)的优化。我们在多个任务和不同难度目标下验证了该方法的有效性,成功构建两个新基准并扩展了流行的$τ$-bench。实验表明,所生成基准的难度平均偏差为5.3%至13.2%,相较基线提升2至4倍。

原文摘要 · Abstract (English)

The rapid progress and widespread deployment of LLMs and LLM-powered agents has outpaced our ability to evaluate them. Hand-crafted, static benchmarks are the primary tool for assessing model capabilities, but these quickly become saturated. In contrast, dynamic benchmarks evolve alongside the models they evaluate, but are expensive to create and continuously update. To address these challenges, we develop BeTaL (Benchmark Tuning with an LLM-in-the-loop), a framework that leverages environment design principles to automate the process of dynamic benchmark design. BeTaL works by parameterizing key design choices in base benchmark templates and uses LLMs to reason through the resulting parameter space to obtain target properties (such as difficulty and realism) in a cost-efficient manner. We validate this approach on its ability to create benchmarks with desired difficulty levels. Using BeTaL, we create two new benchmarks and extend a popular agentic benchmark $τ$-bench. Extensive evaluation on these three tasks and multiple target difficulty levels shows that BeTaL produces benchmarks much closer to the desired difficulty, with average deviations ranging from 5.3% to 13.2% -- a 2-4x improvement over the baselines.

评测基准自动化LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。