用真人完成时间评估AI任务可信度,帮判断AI能否替代人类做复杂工作。
HCAST: Human-Calibrated Autonomy Software Tasks
- 以人类完成时间作为基准,量化AI在189项工程与安全任务中的表现
- 当前AI在人类1小时内完成的任务上成功率70%-80%,4小时以上任务不足20%
- 适合关注AI可靠性、自动化系统评估的研究者和工程师参考
为理解并预测高度自主人工智能系统对社会的影响,需要具备真实世界根基的基准测试,即直接关联AI性能与我们关心的实际影响的度量标准。我们提出HCAST(Human-Calibrated Autonomy Software Tasks),一个包含189项机器学习工程、网络安全、软件工程及通用推理任务的基准。我们从相关领域专家处收集了563个真人基线数据(总计超过1500小时),这些专家在与AI代理相同的条件下执行任务,由此估算出HCAST任务的人类耗时介于1分钟至8小时以上。通过测量任务对人类所需时间,提供了一种直观的评估指标,有助于回答“能否信任一个代理完成需人类花费X小时的任务?”这一问题。我们评估了基于前沿基础模型构建的AI代理在这些任务上的成功概率,发现当前代理在人类耗时少于1小时的任务中成功率可达70%-80%,而在人类耗时超过4小时的任务中成功率低于20%。
原文摘要 · Abstract (English)
To understand and predict the societal impacts of highly autonomous AI systems, we need benchmarks with grounding, i.e., metrics that directly connect AI performance to real-world effects we care about. We present HCAST (Human-Calibrated Autonomy Software Tasks), a benchmark of 189 machine learning engineering, cybersecurity, software engineering, and general reasoning tasks. We collect 563 human baselines (totaling over 1500 hours) from people skilled in these domains, working under identical conditions as AI agents, which lets us estimate that HCAST tasks take humans between one minute and 8+ hours. Measuring the time tasks take for humans provides an intuitive metric for evaluating AI capabilities, helping answer the question "can an agent be trusted to complete a task that would take a human X hours?" We evaluate the success rates of AI agents built on frontier foundation models, and we find that current agents succeed 70-80% of the time on tasks that take humans less than one hour, and less than 20% of the time on tasks that take humans more than 4 hours.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。