用视觉推理任务研究人类如何从少量例子中抽象规则并解决问题。
Exploring Human Behavior During Abstract Rule Inference and Problem Solving with the Cognitive Abstraction and Reasoning Corpus
- 设计了适配人类的抽象推理数据集CogARC,记录高精度行为轨迹。
- 平均准确率约90%(实验1)和80%(实验2),难题耗时更长、策略差异大。
- 即使出错也常收敛,揭示人类在不确定中调整策略的认知机制。
人类在抽象推理中表现出惊人灵活性,能快速从稀疏样本中学习并应用规则。为探究这一能力背后的认知策略,我们引入认知抽象与推理语料库(CogARC),它是原始用于衡量人工智能抽象推理能力的抽象与推理语料库(ARC)的一个适配人类的多样化子集。在两项实验中,共260名参与者自由解答了75个抽象视觉推理问题。解题需从少量示例中推断输入-输出规则,将测试输入转化为正确输出。参与者行为以高时间分辨率记录,包括示例查看、编辑序列和多次提交。总体表现良好:实验1(n=40)平均准确率约90%,实验2(n=220)约80%。但不同问题和参与者间表现差异显著;难题导致更长思考时间,策略分歧更大。任务过程中,参与者响应速度加快,但准确率略有下降,表明对任务结构更熟悉,而非规则学习能力提升。重要的是,即使错误解也常高度收敛,尽管解题路径长短和流畅度各异。部分路径直接高效抵达稳定结果,另一些则经历长时间探索或部分重来后才收敛。这些发现凸显CogARC作为研究人类抽象推理行为的丰富环境,揭示人们如何在不确定性下泛化、误泛化及调整策略。
原文摘要 · Abstract (English)
Humans exhibit remarkable flexibility in abstract reasoning, and can rapidly learn and apply rules from sparse examples. To investigate the cognitive strategies underlying this ability, we introduce the Cognitive Abstraction and Reasoning Corpus (CogARC), a diverse human-adapted subset of the Abstraction and Reasoning Corpus (ARC) which was originally developed to benchmark abstract reasoning in artificial intelligence. Across two experiments, CogARC was administered to a total of 260 human participants who freely generated solutions to 75 abstract visual reasoning problems. Success required inferring input-output rules from a small number of examples to transform the test input into one correct test output. Participants' behavior was recorded at high temporal resolution, including example viewing, edit sequences, and multi-attempt submissions. Participants were generally successful (mean accuracy ~90% for experiment 1 (n=40), ~80% for experiment 2 (n=220) across problems), but performance varied widely across problems and participants. Harder problems elicited longer deliberation times and greater divergence in solution strategies. Over the course of the task, participants initiated responses more quickly but showed a slight decline in accuracy, suggesting increased familiarity with the task structure rather than improved rule-learning ability. Importantly, even incorrect solutions were often highly convergent, even when the problem-solving trajectories differed in length and smoothness. Some trajectories progressed directly and efficiently toward a stable outcome, whereas others involved extended exploration or partial restarts before converging. Together, these findings highlight CogARC as a rich behavioral environment for studying human abstract reasoning, providing insight into how people generalize, misgeneralize, and adapt their strategies under uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。