arXiv:2505.05115cs.AI2025-05被引 5

发现AI代理任务成功率随时间呈指数衰减,可用‘半衰期’描述其性能下降规律。

Is there a half-life for the success rates of AI agents?

  • 用每分钟固定失败率模型解释长任务成功率下降现象。
  • 任务时长越长,成功率按指数衰减,不同代理有各自半衰期。
  • 适用于分析复杂任务中子任务失败的累积效应,适合研究者参考。

基于Kwa等人(2025)的实证工作,本文表明在他们设计的一系列研究-工程任务中,AI代理在长时间任务上的表现可用一个极简数学模型解释——即人类完成该任务每分钟内存在恒定的失败概率。这表明任务成功率随长度呈指数下降,每个代理可被其自身的“半衰期”所刻画。该经验规律使我们能估算代理在不同任务长度下的成功率。模型与数据高度吻合,暗示长期任务失败的深层原因:任务包含越来越多的子任务,任一子任务失败即导致整体失败。该模型是否适用于其他任务集尚不明确,是未来研究的重要方向。

原文摘要 · Abstract (English)

Building on the recent empirical work of Kwa et al. (2025), I show that within their suite of research-engineering tasks the performance of AI agents on longer-duration tasks can be explained by an extremely simple mathematical model -- a constant rate of failing during each minute a human would take to do the task. This implies an exponentially declining success rate with the length of the task and that each agent could be characterised by its own half-life. This empirical regularity allows us to estimate the success rate for an agent at different task lengths. And the fact that this model is a good fit for the data is suggestive of the underlying causes of failure on longer tasks -- that they involve increasingly large sets of subtasks where failing any one fails the task. Whether this model applies more generally on other suites of tasks is unknown and an important subject for further work.

AI代理任务成功率半衰期指数衰减

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。