AI代理常高估任务成功率,预评估比事后复盘更准。
Agentic Uncertainty Reveals Agentic Overconfidence
- 通过任务前后多次预测成功率,测量代理不确定性
- 实际成功仅22%的代理自评成功率高达77%
- 用漏洞查找方式提问能最好校准预测结果
AI代理能否预判自身任务成功率?我们通过在任务执行前、中、后收集成功概率估计来研究代理不确定性。所有结果均显示代理存在过度自信:部分代理实际成功率仅为22%,却自评成功率达77%。反直觉的是,使用较少信息的执行前评估,其判别能力反而优于标准的事后审查,尽管差异并不总显著。对抗性提示将评估重构为漏洞查找任务时,校准效果最佳。
原文摘要 · Abstract (English)
Can AI agents predict whether they will succeed at a task? We study agentic uncertainty by eliciting success probability estimates before, during, and after task execution. All results exhibit agentic overconfidence: some agents that succeed only 22% of the time predict 77% success. Counterintuitively, pre-execution assessment with strictly less information tends to yield better discrimination than standard post-execution review, though differences are not always significant. Adversarial prompting reframing assessment as bug-finding achieves the best calibration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。