用图着色测试大模型的系统性推理能力,发现多数模型在复杂问题上错误率超60%。
Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring
- 以图着色为任务,评估大模型逐步推理与可能性探索能力。
- 所有模型在难问题上错误率均超60%,2-着色4顶点图也未达完美准确率。
- 研究揭示了大模型推理的局限性,适合关注模型可靠性的人阅读。
当前大型语言模型虽具强大问题求解能力,但在推理方面仍存在弱点,相关研究正致力于改进。本文通过图着色任务评估大模型在系统性分步推理和可能空间探索方面的能力,以及语义问题表述的影响。测试对象包括Claude 3.5 Sonnet、Llama 3.1 405B、Gemini 1.5 Pro、GPT-4o、o1-mini 和 DeepSeek-R1,使用 $k$-着色问题数据集($2 \≤ k \≤ 4$,顶点数 $4 \≤ n \≤ 8$),并结合部分算法求解器对问题难度进行分类。结果显示,除 o1-mini 和 R1 外,其余模型在所有表述下困难问题类型上的错误率均超过 60%(o1-mini 超 15%,R1 超 10%),且无模型在 2-着色 4 顶点图这一简单领域实现完全准确。结果凸显大模型在系统推理方面的显著进展及其可靠性限制,尤其与计算成本上升相关。我们预计更复杂的图着色问题及任意复杂度推理问题的程序化生成,将为大模型评测提供尚未开发的潜力。
原文摘要 · Abstract (English)
Contemporary large language models are powerful problem-solving tools, but they exhibit weaknesses in their reasoning abilities which ongoing research seeks to mitigate. We investigate graph coloring as a means of evaluating an LLM's capacities for systematic step-by-step reasoning and possibility space exploration, as well as effects of semantic problem framing. We test Claude 3.5 Sonnet, Llama 3.1 405B, Gemini 1.5 Pro, GPT-4o, o1-mini, and DeepSeek-R1 on a dataset of $k$-coloring problems with $2 \leq k \leq 4$ and vertex count $4 \leq n \leq 8$, using partial algorithmic solvers to further categorize problems by difficulty. In addition to substantial but varying framing effects, we find that all models except o1-mini and R1 exhibit $>60\%$ error rates on difficult problem types in all frames ($>15\%$ for o1-mini and $>10\%$ for R1), and no model achieves perfect accuracy even in the simple domain of 2-coloring 4-vertex graphs. Our results highlight both the considerable recent progress in LLM systematic reasoning and the limits of its reliability, especially in relation to increasing computational costs. We expect that more complex graph coloring problems, and procedural generation of arbitrary-complexity reasoning problems more broadly, offer further untapped potential for LLM benchmarking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。