实测发现AI工具实际效果远低于预期,尤其在开发与医疗领域。
Quantifying the Expectation-Realisation Gap for Agentic AI Systems
- 通过控制实验量化了AI系统在真实场景中的表现差距
- 开发者预估提速24%但实际慢19%,差距达43个百分点
- 适合关注AI落地效果评估、人机协作成本的研究者
具有显著生产率提升预期的自主型AI系统,在实际部署后却表现出系统性差异。我们综述了软件工程、临床文档和临床决策支持领域的控制实验与独立验证,以量化这种预期-实现差距。在软件开发中,经验丰富的开发者预期使用AI工具可提速24%,但实际反而慢了19%,校准误差达43个百分点。在临床文档方面,厂商声称每份病历可节省数分钟,但实测改善不足1分钟,且一款广泛使用的工具未显示统计学显著效果。在临床决策支持中,外部验证性能显著低于开发者报告指标。这些差距主要源于工作流整合摩擦、验证负担、测量标准不匹配,以及受益人群的系统性差异。证据支持采用结构化规划框架,要求明确量化预期收益,并纳入人工监督成本。
原文摘要 · Abstract (English)
Agentic AI systems are deployed with expectations of substantial productivity gains, yet rigorous empirical evidence reveals systematic discrepancies between pre-deployment expectations and post-deployment outcomes. We review controlled trials and independent validations across software engineering, clinical documentation, and clinical decision support to quantify this expectation-realisation gap. In software development, experienced developers expected a 24% speedup from AI tools but were slowed by 19% -- a 43 percentage-point calibration error. In clinical documentation, vendor claims of multi-minute time savings contrast with measured reductions of less than one minute per note, and one widely deployed tool showed no statistically significant effect. In clinical decision support, externally validated performance falls substantially below developer-reported metrics. These shortfalls are driven by workflow integration friction, verification burden, measurement construct mismatches, and systematic variation in who benefits and who does not. The evidence motivates structured planning frameworks that require explicit, quantified benefit expectations with human oversight costs factored in.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。