测试大模型解题能力,发现连有公开解的奥数题也难住它们。
No LLM Solved Yu Tsumura's 554th Problem
- 用一道有公开答案的奥数题检验大模型推理能力
- 所有主流大模型均无法正确解答该题
- 揭示大模型在复杂逻辑推理上的根本缺陷
我们表明,尽管近期大模型在数学竞赛中获得金牌引发乐观情绪,但存在一道问题——宇津村554题——它 a) 在证明复杂度上属于国际数学奥林匹克(IMO)级别,b) 不是导致大模型出错的组合类问题,c) 所需证明技巧少于典型难题,d) 具有公开可得的解法(可能已进入大模型训练数据),e) 仍无法被任何现成的大模型(商业或开源)轻松解决。这表明大模型在高级数学推理方面仍存在系统性短板。
原文摘要 · Abstract (English)
We show, contrary to the optimism about LLM's problem-solving abilities, fueled by the recent gold medals that were attained, that a problem exists -- Yu Tsumura's 554th problem -- that a) is within the scope of an IMO problem in terms of proof sophistication, b) is not a combinatorics problem which has caused issues for LLMs, c) requires fewer proof techniques than typical hard IMO problems, d) has a publicly available solution (likely in the training data of LLMs), and e) that cannot be readily solved by any existing off-the-shelf LLM (commercial or open-source).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。