arXiv:2606.10106cs.SEcs.AI2026-06被引 2

厘清什么是代码生成器的'工具套件',让行业有统一标准。

What makes a harness a harness: necessary and sufficient conditions for an agent harness

论文配图:What makes a harness a harness: necessary and sufficient conditions for an agent harness
图 1 · 摘自论文原文
  • 通过追溯术语演化,定义代理工具套件的必要与充分条件。
  • 提出可操作的判断标准,准确区分工具套件与框架、插件等。
  • 适用于六款真实工具,边界清晰,适合工程与研究参考。

‘代理工具套件’一词在生成式AI软件工程中广泛使用,指将语言模型包装为能作用于代码仓库的编程代理的中间层。当前用法模糊且多义:有时指完整产品(如Claude Code、Codex CLI),有时指评估框架(如SWE-bench harness),有时与代理框架、SDK、IDE插件或编排器混淆。缺乏一个一致可用的参考定义。本文通过分析持久标识符和一手灰色文献(如官方文档、术语表、工程报告),重构该术语从马具到机器学习评估框架,再到代理工具套件的演进脉络。提出一个构成性定义,明确代理工具套件的必要与充分条件,并将其转化为包含与排除测试。据此划清其与代理框架、代理SDK、IDE插件、评估框架及编排器的边界。将该定义应用于六个真实工具(Claude Code、Codex CLI、Aider、Cline、OpenHands、SWE-agent)及边缘案例,测试结果一致。最后提出按设计张力轴划分的研究议程。贡献在于提供一个可操作的代理工具套件定义,建立共享术语体系,指导工程实践并促进代理系统间的科学比较。

原文摘要 · Abstract (English)

The term agent harness now circulates widely in software engineering with generative artificial intelligence. It names the layer that wraps a language model and turns it into a coding agent able to act on a repository. The usage is loose and polysemous. Sometimes the term denotes the whole product (Claude Code, Codex CLI); sometimes it denotes the evaluation scaffold that runs an agent against tasks (the SWE-bench harness); sometimes it gets conflated with an agent framework, an SDK, an IDE plugin, or an orchestrator. What is missing is a reference definition that works as an instrument, one that includes and excludes cases consistently. We build that definition through a conceptual analysis that combines works with persistent identifiers and primary grey-literature sources, such as official documentation, glossaries, and engineering reports. We reconstruct the genealogy of the term, from the horse's tack to the classic test harness, to the machine-learning evaluation harness, and finally to the agent harness. We then propose a constitutive definition that states the necessary and sufficient conditions for a system to be an agent harness, we operationalize it as an inclusion and exclusion test, and we draw the boundary of the concept against an agent framework, an agent SDK, an IDE plugin, an eval harness, and an orchestrator. We apply the definition to six real harnesses (Claude Code, Codex CLI, Aider, Cline, OpenHands, and SWE-agent) and to deliberate edge cases; the test includes and excludes consistently. We close with a research agenda organized by design tension axes. The contribution is an operational definition of agent harness, with a shared vocabulary, able to guide engineering practice and the scientific comparison of agentic systems.

工具套件代理系统定义

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。