自动生成奥数级数学题,提升大模型推理能力
PromptCoT: Synthesizing Olympiad-level Problems for Mathematical Reasoning in Large Language Models
- 基于数学概念与出题逻辑生成复杂题目,模拟专家设计过程
- 在GSM8K、MATH-500等基准上超越现有方法,性能随数据增长稳定提升
- 适合研究大模型数学推理与自动题目生成的学者使用
大语言模型解决复杂数学问题的能力显著提升,尤其在需要高级推理的任务中。然而,高质量、挑战性足够的奥数级题目稀缺,限制了进一步发展。本文提出PromptCoT,一种自动生成高质奥数级数学题的新方法。该方法基于数学概念和出题背后的逻辑,模拟经验丰富的命题者思维过程。我们提供了理论分析,证明最优出题逻辑应同时最大化给定概念下逻辑生成的概率,以及在逻辑与概念共同条件下题目生成的概率。在GSM8K、MATH-500和AIME2024等标准基准上评估,PromptCoT持续优于现有生成方法。此外,实验表明其具有优异的数据可扩展性,在数据规模增大时仍保持高性能,显著超越基线。代码已开源:https://github.com/zhaoxlpku/PromptCoT。
原文摘要 · Abstract (English)
The ability of large language models to solve complex mathematical problems has progressed significantly, particularly for tasks requiring advanced reasoning. However, the scarcity of sufficiently challenging problems, particularly at the Olympiad level, hinders further advancements. In this work, we introduce PromptCoT, a novel approach for automatically generating high-quality Olympiad-level math problems. The proposed method synthesizes complex problems based on mathematical concepts and the rationale behind problem construction, emulating the thought processes of experienced problem designers. We provide a theoretical analysis demonstrating that an optimal rationale should maximize both the likelihood of rationale generation given the associated concepts and the likelihood of problem generation conditioned on both the rationale and the concepts. Our method is evaluated on standard benchmarks including GSM8K, MATH-500, and AIME2024, where it consistently outperforms existing problem generation methods. Furthermore, we demonstrate that PromptCoT exhibits superior data scalability, consistently maintaining high performance as the dataset size increases, outperforming the baselines. The implementation is available at https://github.com/zhaoxlpku/PromptCoT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。