测试GPT-4o1-preview在数理题上的表现,发现进步明显但仍有短板。
Testing GPT-4-o1-preview on math and science problems: A follow-up study
- 用相同题目集测试新版GPT-4o1-preview
- 整体性能提升显著,但空间推理仍弱
- 适合关注大模型数理能力演进的研究者
2023年8月,Scott Aaronson与我报告了使用Wolfram Alpha和Code Interpreter插件对GPT-4在105道高中及大学水平数理题上的测试结果(Davis and Aaronson, 2023)。2024年9月,我使用新发布的GPT-4o1-preview在同一题集上进行了测试。总体来看,性能有显著提升,但仍远未达到完美。尤其涉及空间推理的问题常成为瓶颈。
原文摘要 · Abstract (English)
In August 2023, Scott Aaronson and I reported the results of testing GPT4 with the Wolfram Alpha and Code Interpreter plug-ins over a collection of 105 original high-school level and college-level science and math problems (Davis and Aaronson, 2023). In September 2024, I tested the recently released model GPT-4o1-preview on the same collection. Overall I found that performance had significantly improved, but was still considerably short of perfect. In particular, problems that involve spatial reasoning are often stumbling blocks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。