arXiv:2605.22535cs.AI2026-05被引 9

自动构建真实终端任务基准,评估智能体在复杂操作中的表现

TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks

论文配图:TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks
图 1 · 摘自论文原文
  • 从8万+真实终端记录中自动生成高保真评估任务
  • 顶尖模型在200个验证任务上最高通过率仅62.5%
  • 适合关注真实开发场景下智能体能力的研究者

我们提出TerminalWorld,一个可扩展的数据引擎,能从真实终端操作记录中自动逆向生成高保真评估任务。处理80,870条终端记录后,生成包含1,530个经验证任务的完整基准,覆盖18类真实世界任务,从日常短操作到超过50步的工作流,涵盖1,280个唯一命令。从中筛选出200个人工审核的代表性任务组成Verified子集。在该子集上对八个前沿模型和六个智能体进行综合评估发现,当前系统仍难以应对真实终端工作流,最高通过率仅为62.5%。此外,TerminalWorld捕捉到与现有专家标注基准(如Terminal-Bench)显著不同的真实能力,二者得分相关性极弱(皮尔逊相关系数r=0.20)。该自动化引擎使基准天然具备真实性与可扩展性,可随开发者实践演进而持续评估智能体。数据与代码已公开于https://github.com/EuniAI/TerminalWorld。

原文摘要 · Abstract (English)

We introduce TerminalWorld, a scalable data engine that automatically reverse-engineers high-fidelity evaluation tasks from "in-the-wild" terminal recordings. Processing 80,870 terminal recordings, the engine yields a full benchmark of 1,530 validated tasks, spanning 18 real-world categories, ranging from short everyday operations to workflows exceeding 50 steps, and covering 1,280 unique commands. From these, we curate a Verified subset of 200 representative, manually reviewed tasks. Comprehensive benchmarking on TerminalWorld-Verified across eight frontier models and six agents reveals that current systems still struggle with authentic terminal workflows, achieving a maximum pass rate of only 62.5%. Moreover, TerminalWorld captures real-world terminal capabilities distinct from existing expert-curated benchmarks (e.g., Terminal-Bench), with only a weak correlation to their scores (Pearson r=0.20). The automated engine makes TerminalWorld authentic and scalable by construction, enabling it to evaluate agents in real-world terminal environments as developer practices evolve. Data and code are available at https://github.com/EuniAI/TerminalWorld.

智能体评估终端任务真实场景自动化构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。