如何设计靠谱的终端代理评估任务?关键在对抗性、难度与可读性。
What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design
- 将任务设计视为提示编写会误导评估,应聚焦验证模型真实能力。
- 超15%常见任务存在奖励劫持问题,易被模型钻空子通过虚假路径得分。
- 强调概念性难度而非环境复杂度,适合评估者和研究者参考改进基准设计。
终端代理基准已成为衡量大语言模型编程与系统管理能力的主要指标。随着评估环境市场的发展,任务开发压力增大,常缺乏充分的对抗性审查。本文基于超过一年对Terminal Bench任务的贡献与评审经验,提出良好任务设计指南:任务应具备对抗性、困难性和可读性。多数失败模式——如由AI生成的指令、过度具体的规范、琐碎操作、依赖隐含知识的验证逻辑、测试错误目标、可被奖励劫持的环境——均源于将任务创作误作提示编写。我们系统梳理这些陷阱,指出真正的难度在于概念层面而非环境复杂度,并引用最新实证数据表明,超过15%的主流终端代理基准任务存在奖励劫持现象。本指南旨在为基准维护者、任务贡献者及以基准分数为证据的研究者提供参考。
原文摘要 · Abstract (English)
Terminal-agent benchmarks have become a primary signal for measuring the coding and system-administration capabilities of large language models. As the market for evaluation environments grows, so does the pressure to ship tasks quickly, often without thorough adversarial review of the verification logic. This paper is a guideline for writing good benchmark tasks, drawn from over a year of contributing to and reviewing tasks for Terminal Bench. Most people write benchmark tasks the way they write prompts. They shouldn't. A prompt is designed to help the agent succeed; a benchmark is designed to find out if it can. We argue that good tasks are adversarial, difficult, and legible, and that a large class of common failure modes -- AI-generated instructions, over-prescriptive specifications, clerical difficulty, oracle solutions that assume hidden knowledge, tests that validate the wrong things, and reward-hackable environments -- are predictable consequences of treating task authoring as prompt authoring. We catalog these failure modes, argue that real difficulty is conceptual rather than environmental, and discuss recent empirical evidence that over 15% of tasks in popular terminal-agent benchmarks are reward-hackable. We hope this serves as a useful reference for benchmark maintainers, task contributors, and researchers using benchmark scores as evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。