大规模测试人类在ARC任务中的表现,发现平均正确率超76%。
H-ARC: A Robust Estimate of Human Performance on the Abstraction and Reasoning Corpus Benchmark
- 1729人完成400训练+400测试任务,获得更可靠的基准数据
- 人类在训练集平均正确率达76.2%,测试集为64.2%
- 结果表明人类远超当前AI模型,适合研究通用推理能力
Abstraction and Reasoning Corpus(ARC)是一个用于测试人类和机器在分布外泛化能力的视觉程序合成基准。自2019年以来,人工智能方法在该挑战上进展有限。准确比较人类与机器性能对基准有效性至关重要。此前研究仅使用部分任务或变体评估人类表现,估计不够可靠。本文通过让1729名参与者完成原始ARC数据集全部400个训练任务和400个评估任务,获得更稳健的人类表现估计:训练集平均正确率在73.3%至77.2%之间,实测平均为76.2%;评估集在55.9%至68.9%之间,实测平均为64.2%。此外,800个任务中有790个至少一人在三次尝试内解决,说明绝大多数任务对普通网络众包人员可解。尽管数值略低于早期估计,人类表现仍显著优于现有最先进方法。为促进研究,本文公开了包含所有提交记录与操作轨迹的H-ARC数据集。
原文摘要 · Abstract (English)
The Abstraction and Reasoning Corpus (ARC) is a visual program synthesis benchmark designed to test challenging out-of-distribution generalization in humans and machines. Since 2019, limited progress has been observed on the challenge using existing artificial intelligence methods. Comparing human and machine performance is important for the validity of the benchmark. While previous work explored how well humans can solve tasks from the ARC benchmark, they either did so using only a subset of tasks from the original dataset, or from variants of ARC, and therefore only provided a tentative estimate of human performance. In this work, we obtain a more robust estimate of human performance by evaluating 1729 humans on the full set of 400 training and 400 evaluation tasks from the original ARC problem set. We estimate that average human performance lies between 73.3% and 77.2% correct with a reported empirical average of 76.2% on the training set, and between 55.9% and 68.9% correct with a reported empirical average of 64.2% on the public evaluation set. However, we also find that 790 out of the 800 tasks were solvable by at least one person in three attempts, suggesting that the vast majority of the publicly available ARC tasks are in principle solvable by typical crowd-workers recruited over the internet. Notably, while these numbers are slightly lower than earlier estimates, human performance still greatly exceeds current state-of-the-art approaches for solving ARC. To facilitate research on ARC, we publicly release our dataset, called H-ARC (human-ARC), which includes all of the submissions and action traces from human participants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。