用AI开发对话式引导系统,发现成本估算易被误导。
Building AI-Intensive Software with AI: Early Results and a Cautionary Tale on Measuring Development Cost

- 三层次成本模型:实际AI支出、人工自报工时、人力替代假设。
- 原始19.4倍成本比因两处计算错误虚高,修正后为9.9倍。
- 提醒研究者警惕隐蔽误差,推动可复现的AI开发成本度量方法。
关于人工智能密集型软件开发真实成本的实证报告仍然稀少,且现有报告极易因未暴露的错误而失准。我们报告一项正在进行的案例研究的早期结果:一个六人学生团队在一个学期内,借助广泛的人工智能辅助,构建了一个完整的对话式引导助手——基于RAG的代码聊天、引导式导览、依赖图谱与技术债务分析。我们采用三层成本模型(实际AI支出、人工自报工时、人力替代假设)进行开发过程追踪,最初报告的成本比为19.4倍。后续核查发现两个独立错误:在固定订阅制下误推每令牌成本,以及以错误区域工资定价人力替代假设,两者共同导致成本比被高估约2倍;修正后结果约为9.9倍。我们以此修正作为一项早期、具普遍意义的发现:这两类错误易于犯出,但在最终数字中难以察觉,可能广泛存在于类似报告中。我们提出迈向更稳健、可复现的AI密集型开发成本度量方法的下一步方向。
原文摘要 · Abstract (English)
Empirical reports on the true cost of AI-intensive software development remain scarce, and the few that exist are easy to get wrong in ways that never surface in the final number. We report early results from an ongoing case study: a six-person student team built a full conversational onboarding assistant -- RAG-based code chat, guided tours, dependency graphs, technical-debt analysis -- over one academic term using pervasive AI assistance. We instrumented development with a three-layer cost model (real AI spend, self-reported human effort, human counterfactual) and initially reported a 19.4x cost ratio. A follow-up pass revealed two independent errors -- inferring per-token cost under a flat-rate subscription, and pricing the counterfactual with the wrong regional labor rates -- that together had inflated the ratio by roughly 2x; the corrected figure is ~9.9x. We present this correction as an early, generalizable finding in its own right: both errors are easy to make, invisible in the final number, and plausibly common in similar reports. We outline next steps toward a more robust, replicable costing methodology for AI-intensive development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。