arXiv:2602.05059cs.AI2026-02被引 1

LLM能辅助理解已知图论问题,但无法解决未解难题。

Evaluating Large Language Models on Solved and Unsolved Problems in Graph Theory: Implications for Computing Education

  • 用八阶段评估法检验LLM在图论问题中的数学推理能力
  • 对已解问题生成正确定义、结构识别与专家验证的证明
  • 对未解问题保持诚实不捏造,适合用于概念探索教学

大型语言模型在计算机科学进阶内容(如图论)学习中日益普及。本研究通过八阶段评估流程,考察其在两个相关图论问题上的表现:一个已解决的问题(关于线图优美性),以及一个尚未解决的开放问题。对于已解问题,模型准确给出定义,识别关键结构,无幻觉地回忆相关定理,并构建出经图论专家验证的有效证明。对于开放问题,模型虽能提出连贯的解释和合理的探索策略,但未能推进至解决方案,且未虚构结论,而是明确表达不确定性。这符合提示指令要求,避免编造定理或无根据断言。结果表明,LLM可有效支持已有知识的探索,但在需要新洞察或关键结构性推理的任务中仍存局限。对计算教育而言,应引导学生利用LLM进行概念探究,而正式解题仍需独立验证与严谨论证。

原文摘要 · Abstract (English)

Large Language Models are increasingly used by students to explore advanced material in computer science, including graph theory. As these tools become integrated into undergraduate and graduate coursework, it is important to understand how reliably they support mathematically rigorous thinking. This study examines the performance of a LLM on two related graph theoretic problems: a solved problem concerning the gracefulness of line graphs and an open problem for which no solution is currently known. We use an eight stage evaluation protocol that reflects authentic mathematical inquiry, including interpretation, exploration, strategy formation, and proof construction. The model performed strongly on the solved problem, producing correct definitions, identifying relevant structures, recalling appropriate results without hallucination, and constructing a valid proof confirmed by a graph theory expert. For the open problem, the model generated coherent interpretations and plausible exploratory strategies but did not advance toward a solution. It did not fabricate results and instead acknowledged uncertainty, which is consistent with the explicit prompting instructions that directed the model to avoid inventing theorems or unsupported claims. These findings indicate that LLMs can support exploration of established material but remain limited in tasks requiring novel mathematical insight or critical structural reasoning. For computing education, this distinction highlights the importance of guiding students to use LLMs for conceptual exploration while relying on independent verification and rigorous argumentation for formal problem solving.

图论LLM评估计算教育

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。