为AI编程代理构建六维流程分类体系,揭示其协作机制与短板。
From Prompt to Process: a Process Taxonomy and Comparative Assessment of Frameworks Supporting AI Software Development Agents
- 提出六维流程框架:规范、上下文、角色、执行、验证、可迁移性
- 发现现有框架普遍缺维度,深度与跨平台兼容性难以兼顾
- 适合研究AI开发流程或评估编程代理的学者与开发者
AI编程工具已超越自动补全与聊天助手,演变为具备流程、角色、产出物和验证机制的开发框架。尽管已有研究综述了软件工程中的智能体与大模型,但缺乏对将这些能力转化为实际工作流程的框架的系统分析。我们通过定向检索原始文献,基于功能纳入标准与影响力指标,筛选出六个框架:GitHub Spec Kit、OpenSpec、BMAD Method、Get Shit Done(GSD)、Spec Kitty 和 Reversa。它们分别采用不同路径实现AI驱动开发:完整与轻量级的规范驱动开发、代理驱动的敏捷规划、上下文工程、工作树隔离与评审、以及从遗留系统中恢复操作规范。本文核心贡献是一个六维流程分类体系:规范、上下文、角色、执行、验证、可迁移性,并配套评分标准,使其可复现。应用于六个框架及一个外部案例Spec-Flow,结果表明:已采用流程的框架出现趋同——提示词中心地位下降,持久化产出物、工作契约、可追溯性和人工评审成为降低模糊性与协调代理的关键机制。同时,无一框架全面覆盖所有维度,揭示出流程深度与跨代理可移植性之间的结构性权衡。此外,还识别出反复出现的风险:规范与代码间的漂移、对生成产物的过度信任、社区扩展的脆弱性、平台依赖性以及完整流程缺乏基准评测。最后,提出以中间质量度量、上下文治理、安装安全与可复现性为重点的研究议程。
原文摘要 · Abstract (English)
AI tools for programming are no longer just autocomplete or chat assistants: they organize themselves as development frameworks, with process, roles, artifacts and verification. Recent surveys map agents and LLMs for software engineering, but a study centered on the operational frameworks that turn these capabilities into process is missing. We ran a directed search of primary sources, with a functional inclusion criterion and traction measurement, and selected six frameworks: GitHub Spec Kit, OpenSpec, BMAD Method, Get Shit Done (GSD), Spec Kitty and Reversa. Each attacks AI development through a different path: spec-driven development in full and lightweight variants, agent-driven agile planning, context engineering over the agent, worktree isolation and review, and recovery of operational specifications from legacy systems. Our central contribution is a six-dimension process taxonomy: specification, context, roles, execution, validation and portability, with a scoring rubric that turns it into a replicable instrument. We apply it to the six frameworks and an out-of-sample case, Spec-Flow. Two results stand out. Among frameworks that already adopt some process there is convergence: the isolated prompt loses centrality, and persistent artifacts, work contracts, traceability and human review become mechanisms that reduce ambiguity and coordinate agents. And no framework strongly covers all six dimensions, exposing a structural trade-off between process depth and portability across agents. We also found recurring risks: drift between specification and code, excessive trust in generated artifacts, fragility of community extensions, platform dependence and a lack of benchmarks for the complete process. We close with a research agenda for empirical evaluation, focused on intermediate-quality metrics, context governance, installation security and reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。