用机器评估任务测试人类流体智力,发现其有效且与推理能力强相关。
Bringing Back Rule Induction to Fluid Intelligence Research? An Initial Validation of the ARC-AGI Benchmark in Humans
- 用ARC-AGI任务测量人类规则归纳能力,验证其心理计量特性。
- 100人样本中,该任务与图形推理测验相关性达0.63,显著正相关。
- 首次将机器智能基准用于人类认知研究,推动跨学科协作。
关于流体智力(gf)的两种竞争观点认为:表现受限于工作记忆容量或新关系归纳能力。当前测量主要依赖有限重复规则,而后者虽常被定义却少被测量。ARC-AGI基准以规则归纳为核心,被提议作为人类与人工智能的流体智力衡量工具,但其在人类中的心理计量属性尚未检验。本研究对100名参与者进行首次探索,结果显示该任务组合具有良好的心理计量特性,并与图形推理测验呈显著正相关(ρ = .63)。与图形原创性关联较弱。结果初步支持了ARC-AGI作为人类流体智力测量工具的有效性。未来研究应增加更多规则归纳任务及多变量协变量。本研究独特之处在于首次在人类中测试原本为机器设计的任务,建议系统性地将人工智能基准嵌入人类认知能力的理论网络,以促进更系统的评估与跨学科合作。
原文摘要 · Abstract (English)
Two competing perspectives on fluid intelligence (gf) measures propose that performance is primarily constrained either by working memory capacity or by the ability to induce novel relations. The first perspective is currently dominant in measurement, as evident from the use of a limited set of recurring rules, whereas the second perspective is reflected in many definitions but rarely present in measurement. The ARC-AGI benchmark predominantly requires rule induction and was proposed as a measure of gf for both humans and artificial systems. However, its psychometric properties have not yet been examined in human samples. We therefore investigated the psychometric characteristics and nomological network of ARC-AGI in a first study with 100 participants. A compilation of ARC-AGI items showed good psychometric properties and correlated substantially with figural fluid intelligence as measured by a figural reasoning test ($ρ$ = .63). Associations with figural originality were weak. These findings provide initial support for the validity of ARC-AGI as a measure of human fluid intelligence. Future research should include more rule induction tasks as well as additional multivariate covariates. This study is unusual by studying a task in humans that was initially designed for machines. We suggest systematically embedding AI benchmarks into the nomological network of human cognitive abilities to enable more systematic evaluation and interdisciplinary cooperation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。