arXiv:2409.16125cs.AI2024-09中稿 · ed被引 5

提出两种评估AI能力的新方法,发现它们存在系统性偏差。

Analyzing Probabilistic Methods for Evaluating Agent Capabilities

  • 将任务拆解为子任务或用人作为代理,改进成功率估计
  • 实验显示两种方法均低估真实解决率,尤其专家-最佳-N法更严重
  • 建议结合蒙特卡洛估计理论提升对困难任务的评估准确性

为降低人工智能系统的风险,需准确评估其能力,尤其在能力仅罕见展现时更为困难。Phuong等人提出了两种方法以改善对AI智能体完成特定任务概率的估计:里程碑法将任务分解为子任务,旨在提升整体成功概率估计;专家最佳-N法则利用人类指导作为模型独立表现的代理。我们将其分析为蒙特卡洛估计器,发现两者虽有效降低方差,但引入偏差。实验表明,里程碑法因约束性假设而低估多数真实任务解决率;专家最佳-N法在所有任务上均出现更严重的低估,归因于其固有的重加权因子缺陷。为提高对困难任务中AI能力估计的准确性,建议未来工作应借鉴蒙特卡洛估计器的丰富文献。

原文摘要 · Abstract (English)

To mitigate risks from AI systems, we need to assess their capabilities accurately. This is especially difficult in cases where capabilities are only rarely displayed. Phuong et al. propose two methods that aim to obtain better estimates of the probability of an AI agent successfully completing a given task. The milestone method decomposes tasks into subtasks, aiming to improve overall success rate estimation, while the expert best-of-N method leverages human guidance as a proxy for the model's independent performance. Our analysis of these methods as Monte Carlo estimators reveals that while both effectively reduce variance compared to naive Monte Carlo sampling, they also introduce bias. Experimental results demonstrate that the milestone method underestimates true solve rates for many real-world tasks due to its constraining assumptions. The expert best-of-N method exhibits even more severe underestimation across all tasks, attributed to an inherently flawed re-weighting factor. To enhance the accuracy of capability estimates of AI agents on difficult tasks, we suggest future work should leverage the rich literature on Monte Carlo Estimators.

AI评估蒙特卡洛能力评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。