arXiv:2410.07114cs.CYcs.AI2024-10被引 29

o1-preview模型在数学考试中接近满分,展现深度推理能力。

System 2 thinking in OpenAI's o1-preview model: Near-perfect performance on a mathematics exam

  • 通过重复提示并取多数答案,提升推理准确率。
  • 两次考试分别得76和74分,远超GPT-4o的66和62分。
  • 在新考题上仍表现优异,排除了数据泄露可能。

人类认知常分为快速直觉的System 1与慢速分析的System 2。此前大语言模型被认为缺乏System 2式的深度推理能力。2024年9月,OpenAI推出o1系列模型,旨在实现此类推理。本研究对o1-preview模型两次测试荷兰'数学B'高考卷,得分分别为76和74/76。对比之下,仅有24名荷兰考生(共16,414人)获满分。而GPT-4o得分66和62,高于荷兰学生平均分40.63。两模型均未接触试卷图表。为规避模型污染风险(因o1-preview与GPT-4o知识截止晚于试题公开),我们使用一份截止日期后的全新数学考试复测,结果表明o1-preview仍达97.8百分位,证明污染非主因。同时发现o1-preview输出存在波动,偶有正确或错误结果,采用自一致性策略(多次生成取多数答案)可有效识别正确解。结论:o1系列潜力巨大,但需关注其不确定性。

原文摘要 · Abstract (English)

The processes underlying human cognition are often divided into System 1, which involves fast, intuitive thinking, and System 2, which involves slow, deliberate reasoning. Previously, large language models were criticized for lacking the deeper, more analytical capabilities of System 2. In September 2024, OpenAI introduced the o1 model series, designed to handle System 2-like reasoning. While OpenAI's benchmarks are promising, independent validation is still needed. In this study, we tested the o1-preview model twice on the Dutch 'Mathematics B' final exam. It scored a near-perfect 76 and 74 out of 76 points. For context, only 24 out of 16,414 students in the Netherlands achieved a perfect score. By comparison, the GPT-4o model scored 66 and 62 out of 76, well above the Dutch students' average of 40.63 points. Neither model had access to the exam figures. Since there was a risk of model contami-nation (i.e., the knowledge cutoff for o1-preview and GPT-4o was after the exam was published online), we repeated the procedure with a new Mathematics B exam that was published after the cutoff date. The results again indicated that o1-preview performed strongly (97.8th percentile), which suggests that contamination was not a factor. We also show that there is some variability in the output of o1-preview, which means that sometimes there is 'luck' (the answer is correct) or 'bad luck' (the output has diverged into something that is incorrect). We demonstrate that the self-consistency approach, where repeated prompts are given and the most common answer is selected, is a useful strategy for identifying the correct answer. It is concluded that while OpenAI's new model series holds great potential, certain risks must be considered.

推理模型数学考试自一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。