评测大模型代理中执行层配置对实际工作流的影响。
Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows

- 构建106个真实代理任务,对比不同执行配置的效果。
- 5194次运行显示模型与执行层组合差异显著。
- 适合研究可靠、可审计代理系统的设计者。
大型语言模型代理正被部署为使用工具、修改工作空间并生成具体成果的可执行系统。此类工作流的表现不仅取决于基础模型,还依赖于执行层(harness)——负责管理上下文、工具、状态、约束、权限、追踪和恢复的系统层。然而现有基准通常抽象掉执行过程,或固定执行层,难以研究执行层差异。本文提出Harness-Bench,一个诊断性基准,用于评估在真实代理工作流中不同配置层级的执行层影响。该基准在共享任务环境、预算和评估协议下,对多个模型后端的代表性执行配置进行评估,同时保留各执行层的原生执行行为。基准包含106个沙箱离线任务,基于实际代理使用模式构建,并经人工审核确保现实性、可解性、可验证性和完整性。每次运行记录最终成果、执行轨迹、使用统计和验证器输出,支持超越最终完成度的分析。在5,194条执行轨迹中,观察到模型-执行层组合在完成率、过程质量、效率和失败行为上存在显著差异。结果表明,代理能力应报告为模型-执行层配置级别的性能,而非仅归因于基础模型。分析进一步识别出重复出现的执行对齐失败,即合理推理与工具反馈、工作区状态、证据或可验证输出契约脱节。Harness-Bench为诊断和改进可靠、高效、可审计的代理执行栈提供了可复现的基础。
原文摘要 · Abstract (English)
LLM agents are increasingly deployed as executable systems that use tools, modify workspaces, and produce concrete artifacts. In such workflows, performance depends not only on the base model, but also on the harness: the system layer that manages context, tools, state, constraints, permissions, tracing, and recovery. However, existing benchmarks typically abstract away execution, compare complete agent systems, or hold the harness fixed, making execution-layer variation difficult to study. We introduce Harness-Bench, a diagnostic benchmark for evaluating configuration-level harness effects in realistic agent workflows. Harness-Bench evaluates representative harness configurations across multiple model backends under shared task environments, budgets, and evaluation protocols, while preserving each harness's native execution behavior. The benchmark contains 106 sandboxed offline tasks constructed from practical agent-use patterns and manually reviewed for realism, solvability, oracle-checkability, and integrity. Each run records final artifacts, execution traces, usage statistics, and validator outputs, enabling analysis beyond final completion. Across 5,194 execution trajectories, we observe substantial variation in completion, process quality, efficiency, and failure behavior across model-harness pairings. These results suggest that agent capability should be reported at the model-harness configuration level rather than attributed to the base model alone. Our analysis further identifies recurring execution-alignment failures, where plausible reasoning becomes decoupled from tool feedback, workspace state, evidence, or verifiable output contracts. Harness-Bench provides a reproducible foundation for diagnosing and improving reliable, efficient, and auditable agent execution stacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。