测试大模型能否提出数学新解法,发现部分模型有创造力
Assessing the Creativity of LLMs in Proposing Novel Solutions to Mathematical Problems
- 设计新基准CreativeMath,评估模型在已知解法后提新解的能力
- Gemini-1.5-Pro在生成新颖解法上表现最优,但整体创造力差异大
- 适合关注AI创新力、数学辅助发现的研究者和开发者
人工智能的数学能力复杂且多维。现有研究多关注AI解题的正确性,而本文认为,除了给出正确答案,AI还应能或协助人类提出数学问题的新解法。本研究探索大语言模型(LLMs)在数学推理中的创造性潜力,这一方向此前关注较少。我们提出全新框架与基准CreativeMath,涵盖从中学课程到奥数级别的题目,用于评估模型在已有解法基础上提出创新解法的能力。实验表明,尽管LLMs在常规数学任务中表现良好,其创造性解决问题的能力差异显著。值得注意的是,Gemini-1.5-Pro在生成新颖解法方面优于其他模型。该研究开辟了评估AI创造力的新方向,揭示了大模型在推动数学创新中的优势与局限,为未来AI辅助数学发现奠定基础。
原文摘要 · Abstract (English)
The mathematical capabilities of AI systems are complex and multifaceted. Most existing research has predominantly focused on the correctness of AI-generated solutions to mathematical problems. In this work, we argue that beyond producing correct answers, AI systems should also be capable of, or assist humans in, developing novel solutions to mathematical challenges. This study explores the creative potential of Large Language Models (LLMs) in mathematical reasoning, an aspect that has received limited attention in prior research. We introduce a novel framework and benchmark, CreativeMath, which encompasses problems ranging from middle school curricula to Olympic-level competitions, designed to assess LLMs' ability to propose innovative solutions after some known solutions have been provided. Our experiments demonstrate that, while LLMs perform well on standard mathematical tasks, their capacity for creative problem-solving varies considerably. Notably, the Gemini-1.5-Pro model outperformed other LLMs in generating novel solutions. This research opens a new frontier in evaluating AI creativity, shedding light on both the strengths and limitations of LLMs in fostering mathematical innovation, and setting the stage for future developments in AI-assisted mathematical discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。