构建运行时框架,让大模型编程代理更可靠。
AI Harness Engineering: A Runtime Substrate for Foundation-Model Software Agents
- 提出由运行时支架管理模型观察、行动与反馈的系统
- 实测显示高阶支架可生成可验证的修复日志与报告
- 适合研究智能编程代理底层系统的设计者
基础模型已改变自动化代码生成,但自主软件工程代理在真实开发环境中仍不可靠。现有解释归因于模型能力不足,我们提出新视角:软件工程能力源于模型-支架-环境协同系统,其中运行时支架(harness)负责管理模型对项目的观察、操作、反馈接收及变更完成判定。我们形式化该支架为人工智能支架工程,识别出十一项核心职责:任务定义、上下文选择、工具访问、项目记忆、任务状态、可观测性、失败归因、验证、权限控制、熵审计与干预记录。通过四级阶梯(H0-H3)逐步暴露运行时支持,并设计基于追踪的评估协议,将每次代理运行转化为可审计的事件包。在受控验证任务中,低阶支架仅输出最终补丁,高阶支架则产出可复现日志、失败归因、确定性需求检查与结构化验证报告。该框架将自主软件工程的核心问题,从‘模型能否生成补丁’转变为‘系统能否生成可验证、可归因、可维护的变更’。我们提出了未来基础模型编程代理所需运行时系统的研究路线。
原文摘要 · Abstract (English)
Foundation models have transformed automated code generation, yet autonomous software-engineering agents remain unreliable in realistic development settings. The dominant explanation locates this gap in model capability. We propose a different locus: software-engineering capability emerges from a model-harness-environment system, in which a runtime substrate -- the harness -- mediates how a foundation-model agent observes a project, acts on it, receives feedback, and establishes that a change is complete. We formalize this substrate as an AI Harness Engineering and identify eleven component responsibilities: task specification, context selection, tool access, project memory, task state, observability, failure attribution, verification, permissions, entropy auditing, and intervention recording. We operationalize the harness through a four-level ladder (H0-H3) that progressively exposes runtime support to the agent, and we propose a trace-based evaluation protocol that converts each agent run into an auditable episode package. Applied to a controlled validation task, the framework yields episode packages whose evidence structure varies systematically with harness level: lower levels produce only a final patch, higher levels produce reproduction logs, failure attributions, deterministic requirement checks, and structured verification reports. The framework reframes the central question of autonomous software engineering from whether a foundation model can produce a patch to whether the model-harness-environment system can produce a verifiably correct, attributed, and maintainable change. We outline a research program for the runtime systems that foundation-model software agents will require.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。