arXiv:2412.05753cs.CYcs.AI2024-12被引 5

OpenAI o1在高阶思维任务中多项超越人类,尤其在逻辑与科学推理上表现突出。

Can OpenAI o1 outperform humans in higher-order cognitive thinking?

  • 采用多维度认知测试评估o1-preview模型表现
  • 在批判性写作、系统思维等6项任务中均优于人类平均水平
  • 适合关注AI辅助教育与认知能力评估的研究者参考

本研究评估了OpenAI o1-preview模型在批判性思维、系统思维、计算思维、数据素养、创造性思维、逻辑推理和科学推理等高阶认知领域的表现。通过多个基准测试,将其与不同教育水平的人类参与者进行对比。o1-preview在Ennis-Weir批判性思维作文测试(EWCTET)中平均得分为24.33,显著高于本科生(13.8)和研究生(18.39)(z值分别为1.60和0.90)。在系统思维测试(Lake Urmia Vignette)中得分46.1(SD=4.12),远超人类均值20.08(SD=8.13,z=3.20)。数据素养维度上得分为8.60(SD=0.70),高于人类后测均值4.17(SD=2.02,z=2.19)。创造性思维原始性得分2.98(SD=0.73),高于人类均值1.74(z=0.71)。逻辑推理(LogiQA)准确率达90%(SD=10%),高于人类的86%(SD=6.5%,z=0.62)。科学推理任务中平均得分0.99(SD=0.12),接近完美,超过人类最高分0.85(SD=0.13,z=1.78)。尽管在结构化任务中表现优异,但其在问题解决与适应性推理方面仍存在局限。结果表明AI可辅助教育评估,但需加强伦理监管与能力拓展。

原文摘要 · Abstract (English)

This study evaluates the performance of OpenAI's o1-preview model in higher-order cognitive domains, including critical thinking, systematic thinking, computational thinking, data literacy, creative thinking, logical reasoning, and scientific reasoning. Using established benchmarks, we compared the o1-preview models's performance to human participants from diverse educational levels. o1-preview achieved a mean score of 24.33 on the Ennis-Weir Critical Thinking Essay Test (EWCTET), surpassing undergraduate (13.8) and postgraduate (18.39) participants (z = 1.60 and 0.90, respectively). In systematic thinking, it scored 46.1, SD = 4.12 on the Lake Urmia Vignette, significantly outperforming the human mean (20.08, SD = 8.13, z = 3.20). For data literacy, o1-preview scored 8.60, SD = 0.70 on Merk et al.'s "Use Data" dimension, compared to the human post-test mean of 4.17, SD = 2.02 (z = 2.19). On creative thinking tasks, the model achieved originality scores of 2.98, SD = 0.73, higher than the human mean of 1.74 (z = 0.71). In logical reasoning (LogiQA), it outperformed humans with average 90%, SD = 10% accuracy versus 86%, SD = 6.5% (z = 0.62). For scientific reasoning, it achieved near-perfect performance (mean = 0.99, SD = 0.12) on the TOSLS,, exceeding the highest human scores of 0.85, SD = 0.13 (z = 1.78). While o1-preview excelled in structured tasks, it showed limitations in problem-solving and adaptive reasoning. These results demonstrate the potential of AI to complement education in structured assessments but highlight the need for ethical oversight and refinement for broader applications.

AI认知批判性思维模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。