AI项目用传统估算方法常出错,因五个关键假设在大模型场景下失效。
Five Fatal Assumptions: Why T-Shirt Sizing Systematically Fails for AI Projects
- 提出五种被忽视的估算假设,揭示其在AI项目中的失效机制
- 发现非线性性能跃迁与数据扰动的级联效应,导致进度严重偏差
- 推荐基于决策节点的迭代估算法,更适合复杂AI研发
敏捷估算方法,特别是T恤尺码法,在软件开发中因其简便性广受欢迎。然而,当应用于人工智能项目(尤其是涉及大语言模型和多智能体系统的项目)时,结果往往系统性地误导。本文基于实证分析,指出在使用该方法时常见的五个基础假设:(1)工作量与努力成线性关系,(2)过往经验可重复适用,(3)工作量与时间可互换,(4)任务可分解,(5)完成标准确定。这些假设在传统软件中通常成立,但在AI环境中普遍失效,表现为非线性性能跃迁、复杂交互界面及“强耦合”——数据微小变化会引发整个系统连锁反应。通过分析多智能体系统失败案例、缩放规律及多轮对话的不可靠性,本文提出‘检查点估算’方法:一种以人为中心、分阶段设置决策门限的迭代式估算策略,根据开发过程中的实际发现动态评估范围与可行性,而非依赖初始假设。本文面向负责规划与交付AI项目的工程经理、技术负责人和产品负责人。
原文摘要 · Abstract (English)
Agile estimation techniques, particularly T-shirt sizing, are widely used in software development for their simplicity and utility in scoping work. However, when we apply these methods to artificial intelligence initiatives -- especially those involving large language models (LLMs) and multi-agent systems -- the results can be systematically misleading. This paper shares an evidence-backed analysis of five foundational assumptions we often make during T-shirt sizing. While these assumptions usually hold true for traditional software, they tend to fail in AI contexts: (1) linear effort scaling, (2) repeatability from prior experience, (3) effort-duration fungibility, (4) task decomposability, and (5) deterministic completion criteria. Drawing on recent research into multi-agent system failures, scaling principles, and the inherent unreliability of multi-turn conversations, we show how AI development breaks these rules. We see this through non-linear performance jumps, complex interaction surfaces, and "tight coupling" where a small change in data cascades through the entire stack. To help teams navigate this, we propose Checkpoint Sizing: a more human-centric, iterative approach that uses explicit decision gates where scope and feasibility are reassessed based on what we learn during development, rather than what we assumed at the start. This paper is intended for engineering managers, technical leads, and product owners responsible for planning and delivering AI initiatives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。