arXiv:2607.08964cs.AI2026-07被引 8

新基准测试长时程任务,让智能体在复杂流程中获得过程奖励。

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

论文配图:Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
图 1 · 摘自论文原文
  • 设计46个长周期任务,细分为可评分子任务,实现过程奖励
  • 平均需990万token、85分钟完成,强模型仅15.2%达标
  • 适合评估长期规划与迭代调试能力的智能体研究

AI智能体已能自主完成短时明确任务,但现有终端基准多聚焦几分钟内完成的简单问题,仅以最终结果评判,忽略中间进展与部分解法,导致奖励信号稀疏,无法全面反映智能体能力。本文提出Long-Horizon-Terminal-Bench,包含46个长周期任务,覆盖实验复现、软件工程、多模态分析、交互游戏和科学计算等九类。每个任务采用参考解或仿真引擎,进一步分解为细粒度可评分子任务,支持密集中间奖励与部分得分,可评估智能体在开放流程中的进展程度。任务通常需数百次试运行、数分钟至数小时执行,考验长期规划、长上下文管理与迭代调试能力。评估15个前沿模型发现,平均每任务消耗990万token,约231次试运行,耗时85.3分钟;最强模型在部分奖励阈值0.95下通过率15.2%,完美奖励阈值1.0下为10.9%,模型均值分别为4.3%与1.7%。结果揭示显著提升空间。进一步分析失败模式并发布基准,助力未来长周期终端智能体研究。

原文摘要 · Abstract (English)

AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome. This setup overlooks intermediate progress and partial solutions, yielding sparse reward signals and an incomplete picture of agent capability. We introduce Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing. Each task follows a Terminal-Bench-style setup with a reference solution or simulation engine, but is further decomposed into fine-grained graded subtasks. This design enables dense intermediate rewards and partial credit, allowing evaluation to capture not only whether an agent reaches the final goal, but also how far it progresses on open-ended workflows. Tasks in Long-Horizon-Terminal-Bench typically require hundreds of episodes and minutes to hours of execution, stressing long-horizon planning, long-context management, and iterative debugging rather than one-shot problem solving. We evaluate 15 frontier models and find that agents consume on average 9.9M tokens per task, with roughly 231 episodes and 85.3 minutes of execution time per run, making Long-Horizon-Terminal-Bench more demanding than prior terminal-based benchmarks. Even the strongest tested model achieves 15.2% pass@1 at a partial-reward threshold of 0.95 and 10.9% at a perfect-reward threshold of 1.0, while the mean pass rate across models is 4.3% and 1.7% under the two thresholds, respectively. These results reveal headroom for improvement. We further analyze failure modes and error patterns, and release Long-Horizon-Terminal-Bench to support future progress on long-horizon terminal agents.

智能体评估长时程任务密集奖励基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。