测试OpenAI-o1在数学推理中的表现,发现其不依赖记忆解题。
OpenAI-o1 AB Testing: Does the o1 model really do good reasoning in math problem solving?
- 用国际奥数和中国集训队题目对比测试模型推理能力
- 两组数据表现无显著差异,说明未明显依赖记忆
- 适合关注大模型真实推理能力的研究者
OpenAI的Orion-1模型声称具备更强的逻辑推理能力,但有观点认为其表现可能源于对训练数据中解题方案的「记忆」,导致在未见问题上表现不佳。本文通过两个数据集进行对比实验:一组为公开的国际数学奥林匹克(IMO)题目,另一组为难度相当但较少公开的中国国家队集训营(CNT)题目。对每道题的模型回答进行标注并比较性能。结果表明,两组数据上的表现无显著差异,未发现模型明显依赖记忆。此外,通过案例分析进一步探讨了模型响应特征。
原文摘要 · Abstract (English)
The Orion-1 model by OpenAI is claimed to have more robust logical reasoning capabilities than previous large language models. However, some suggest the excellence might be partially due to the model "memorizing" solutions, resulting in less satisfactory performance when prompted with problems not in the training data. We conduct a comparison experiment using two datasets: one consisting of International Mathematics Olympiad (IMO) problems, which is easily accessible; the other one consisting of Chinese National Team Training camp (CNT) problems, which have similar difficulty but not as publically accessible. We label the response for each problem and compare the performance between the two datasets. We conclude that there is no significant evidence to show that the model relies on memorizing problems and solutions. Also, we perform case studies to analyze some features of the model's response.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。