arXiv:2410.22340cs.CYcs.AI2024-10被引 6

测试GPT-4o1-preview在数理题上的表现,发现进步明显但仍有短板。

Testing GPT-4-o1-preview on math and science problems: A follow-up study

  • 用相同题目集测试新版GPT-4o1-preview
  • 整体性能提升显著,但空间推理仍弱
  • 适合关注大模型数理能力演进的研究者

2023年8月,Scott Aaronson与我报告了使用Wolfram Alpha和Code Interpreter插件对GPT-4在105道高中及大学水平数理题上的测试结果(Davis and Aaronson, 2023)。2024年9月,我使用新发布的GPT-4o1-preview在同一题集上进行了测试。总体来看,性能有显著提升,但仍远未达到完美。尤其涉及空间推理的问题常成为瓶颈。

原文摘要 · Abstract (English)

In August 2023, Scott Aaronson and I reported the results of testing GPT4 with the Wolfram Alpha and Code Interpreter plug-ins over a collection of 105 original high-school level and college-level science and math problems (Davis and Aaronson, 2023). In September 2024, I tested the recently released model GPT-4o1-preview on the same collection. Overall I found that performance had significantly improved, but was still considerably short of perfect. In particular, problems that involve spatial reasoning are often stumbling blocks.

大模型数学推理GPT-4

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。