arXiv:2506.18880cs.CLcs.AI2025-06NeurIPS被引 49

测试大模型在数学中跳出常规思维的能力,发现其创新解题仍严重不足。

OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

  • 构建三轴评估框架,检验模型在探索、组合与颠覆性推理上的泛化能力
  • 顶尖模型在复杂问题上表现急剧下降,组合与颠覆性推理几乎无进步
  • 适合关注AI数学创造力、模型局限性的研究者与开发者

近期大规模语言模型(如 DeepSeek-R1)在奥数级数学基准上表现优异,但往往依赖有限策略,难以应对需要新思路的问题。为此,我们提出 OMEGA——一个受博登创造性分类启发的受控多样化基准,包含三个外分布泛化维度:(1) 探索性,将已知技能应用于更复杂同领域问题;(2) 组合性,整合此前孤立学习的推理技能以解决新问题;(3) 颠覆性,采用非常规策略突破原有方法。OMEGA 通过模板生成器在几何、数论、代数、组合、逻辑与谜题等领域生成可验证的训练-测试对,解决方案采用符号、数值或图形方式验证。我们评估了前沿大模型,发现随着问题复杂度上升,性能显著下降。对 Qwen 系列模型在所有泛化设置下的微调显示,探索性泛化有明显提升,而组合性泛化依然受限,颠覆性推理几乎无改进。通过精细识别这些失败模式,OMEGA 为推动模型超越机械熟练度、迈向真正数学创造力奠定了基础。

原文摘要 · Abstract (English)

Recent large-scale language models (LLMs) with long Chain-of-Thought reasoning-such as DeepSeek-R1-have achieved impressive results on Olympiad-level mathematics benchmarks. However, they often rely on a narrow set of strategies and struggle with problems that require a novel way of thinking. To systematically investigate these limitations, we introduce OMEGA-Out-of-distribution Math Problems Evaluation with 3 Generalization Axes-a controlled yet diverse benchmark designed to evaluate three axes of out-of-distribution generalization, inspired by Boden's typology of creativity: (1) Exploratory-applying known problem solving skills to more complex instances within the same problem domain; (2) Compositional-combining distinct reasoning skills, previously learned in isolation, to solve novel problems that require integrating these skills in new and coherent ways; and (3) Transformative-adopting novel, often unconventional strategies by moving beyond familiar approaches to solve problems more effectively. OMEGA consists of programmatically generated training-test pairs derived from templated problem generators across geometry, number theory, algebra, combinatorics, logic, and puzzles, with solutions verified using symbolic, numerical, or graphical methods. We evaluate frontier (or top-tier) LLMs and observe sharp performance degradation as problem complexity increases. Moreover, we fine-tune the Qwen-series models across all generalization settings and observe notable improvements in exploratory generalization, while compositional generalization remains limited and transformative reasoning shows little to no improvement. By isolating and quantifying these fine-grained failures, OMEGA lays the groundwork for advancing LLMs toward genuine mathematical creativity beyond mechanical proficiency.

数学推理大模型评估创造性智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。