评估编码智能体不能只看模型,系统各环节都影响可靠性。
Engineering Reliable Coding Agents: Evaluating and Operating the System Around the Model
- 构建系统级评估框架,覆盖执行环境、状态管理、检索等多层
- 发现193项关键实践,其中56项深入开发,可提升系统可靠性
- 适合研究者和工程师构建稳定、可复现的编码智能体系统
AI编码智能体常以模型形式评估,却作为系统部署。其可靠性不仅取决于模型能力,还受执行环境、检索机制、记忆与状态管理、权限控制、审查界面及资源分配等因素影响。本文通过整合164篇学术文献、100份实践记录、29个基准测试和17个作者-系统案例,采用多视角评审、定向更新审计、软件工程覆盖分析与分布式系统证据合成,揭示大量所谓模型失败实源于系统其他环节;单一层级改进常无法传导至端到端效果。提出依赖链视角:任务设计、执行环境、检索、状态管理、验证与可观测性中的缺陷会破坏下游结论。贡献包括:206条可靠性记录(含193项受控实践,56项深度开发,13项研究方向)、证据账本、跨生命周期依赖与修复不对称框架、实际运行系统的度量与故障案例、可运行的评估与可靠性协议,以及五项具备证据图谱的可复用技能。整体提供区分模型能力与基础设施影响的方法,支持可辩护的评估设计与容错系统构建。审查具有结构性而非全面性,证据强度因主题而异,结果依赖工作负载与配置。方法明确记录已执行与未执行的搜索路径,限制证据评级范围。
原文摘要 · Abstract (English)
AI coding agents are commonly evaluated as models but deployed as systems. Their reliability depends not only on model capability, but on the harness, execution state, retrieval, memory and state management, permissions, review interfaces, and resource allocation. This monograph examines those boundaries and develops a framework for evaluating and operating coding agents reliably. It synthesizes 164 scholarly works, 100 practitioner records, 29 benchmark records, and 17 author-system case records through a structured multivocal review, targeted update audits, software-engineering coverage analysis, and distributed-systems evidence synthesis. Across this evidence, many apparent model failures originate elsewhere in the system, while improvements at one layer often fail to propagate to end-to-end outcomes. Evaluation and operation are treated as a dependency chain in which weaknesses in task construction, execution environments, retrieval, state management, verification, or observability can invalidate downstream conclusions. The monograph contributes a versioned catalog of 206 reliability records: 193 gated practices, including 56 developed in depth, plus 13 research leads; an evidence ledger; a framework for dependency and repair asymmetry across the agent lifecycle; measurements and failure cases from operated agent systems; runnable evaluation and reliability protocols; and five reusable agent skills with evidence maps. Together, these provide a system-level methodology for distinguishing model capability from infrastructure effects, designing defensible evaluations, and building systems that recover safely when components fail. The review is structured rather than exhaustive, evidence strength varies by topic, and results depend on workload and configuration. The methods record which search lanes were executed, which remain unexecuted, and limits on evidence-grading claims.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。