arXiv:2601.11868cs.SEcs.AI2026-01被引 388

构建真实终端场景下的高难度智能体评测基准,挑战前沿模型能力边界。

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

  • 基于真实工作流设计89个终端任务,环境与解法均经人工编写。
  • 顶尖模型在该基准上平均得分低于65%,暴露通用能力短板。
  • 开源数据集与评估工具,助力开发者优化智能体系统。

AI智能体有望在未来自主完成多样化领域的长期复杂任务。然而现有基准或缺乏现实性,或难度不足,难以有效评估前沿模型。为此,我们提出Terminal-Bench 2.0:一个精心设计的高难度基准,包含89个源自真实工作流的计算机终端任务。每个任务均配备独立环境、人工编写的解决方案及全面的验证测试。实验表明,当前前沿模型和智能体在该基准上的平均得分低于65%,并通过错误分析揭示了模型与智能体改进的关键方向。我们已将数据集与评估工具公开于https://www.tbench.ai/,以支持后续研究与开发。

原文摘要 · Abstract (English)

AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not sufficiently difficult to meaningfully measure frontier models. To this end, we present Terminal-Bench 2.0: a carefully curated hard benchmark composed of 89 tasks in computer terminal environments inspired by problems from real workflows. Each task features a unique environment, human-written solution, and comprehensive tests for verification. We show that frontier models and agents score less than 65\% on the benchmark and conduct an error analysis to identify areas for model and agent improvement. We publish the dataset and evaluation harness to assist developers and researchers in future work at https://www.tbench.ai/ .

智能体评测终端任务基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。