代码不仅是输出,更是智能体运行的基础设施。
Code as Agent Harness

- 用代码作为智能体的通用操作基础,统一连接推理、行动与环境建模。
- 支持长周期任务规划、记忆与工具调用,实现可验证的自适应执行。
- 适合开发自动化系统、智能编程助手及企业级工作流,推动可执行智能体落地。
近期大语言模型在理解与生成代码方面展现出强大能力,涵盖编程竞赛到项目级软件工程。在新兴的智能体系统中,代码已不仅是目标输出,更逐渐成为智能体推理、行动、环境建模和基于执行的验证的运行基础。本文提出‘代码作为智能体底座’(code as agent harness)的统一视角,聚焦代码在智能体基础设施中的核心地位。围绕三个相互关联的层面展开研究:第一,底座接口层,探讨代码如何连接智能体与推理、动作及环境建模;第二,底座机制层,分析长期任务执行中的规划、记忆、工具使用,以及反馈驱动的控制与优化,提升底座的可靠性与适应性;第三,底座扩展层,讨论从单智能体向多智能体系统的演进,共享代码资产支持多智能体协作、审查与验证。文中总结了代表性方法与实际应用,涵盖编码助手、GUI/OS自动化、具身智能体、科学发现、个性化推荐、DevOps与企业流程。同时指出开放挑战:超越任务最终成功的评估、不完整反馈下的验证、无回归的底座优化、多智能体间一致的共享状态、安全关键动作的人工监管,以及向多模态环境的拓展。通过将代码置于智能体工程的核心,本综述为构建可执行、可验证、有状态的智能体系统提供统一路线图。
原文摘要 · Abstract (English)
Recent large language models (LLMs) have demonstrated strong capabilities in understanding and generating code, from competitive programming to repository-level software engineering. In emerging agentic systems, code is no longer only a target output. It increasingly serves as an operational substrate for agent reasoning, acting, environment modeling, and execution-based verification. We frame this shift through the lens of agent harnesses and introduce code as agent harness: a unified view that centers code as the basis for agent infrastructure. To systematically study this perspective, we organize the survey around three connected layers. First, we study the harness interface, where code connects agents to reasoning, action, and environment modeling. Second, we examine harness mechanisms: planning, memory, and tool use for long-horizon execution, together with feedback-driven control and optimization that make harness reliable and adaptive. Third, we discuss scaling the harness from single-agent systems to multi-agent settings, where shared code artifacts support multi-agent coordination, review, and verification. Across these layers, we summarize representative methods and practical applications of code as agent harness, spanning coding assistants, GUI/OS automation, embodied agents, scientific discovery, personalization and recommendation, DevOps, and enterprise workflows. We further outline open challenges for harness engineering, including evaluation beyond final task success, verification under incomplete feedback, regression-free harness improvement, consistent shared state across multiple agents, human oversight for safety-critical actions, and extensions to multimodal environments. By centering code as the harness of agentic AI, this survey provides a unified roadmap toward executable, verifiable, and stateful AI agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。