arXiv:2607.23513cs.CLcs.AI2026-07

测试发现,图表对大模型逻辑推理帮助有限。

Do Diagrams Help Large Language Models Reason? Evidence from Syllogistic Reasoning

  • 对比自然语言、逻辑符号、线性图和欧拉图四种表示方式
  • 模型在蕴含和矛盾问题上表现良好,但中立问题错误率高
  • 适合研究大模型推理机制与可视化辅助效果的学者

图表广泛用于支持逻辑推理,以往研究显示欧拉图等表示形式可提升人类推理表现。近期工作也探索了其对大语言模型(LLMs)的影响。本文比较了四种表示方式在三段论推理中的效果:自然语言、逻辑符号、线性图和欧拉图。基于Ando等(2024)提供的285道题目,评估了Claude 3.5 Sonnet与GPT-4o-mini两个主流模型。结果表明,图表表示并未一致提升性能。尽管模型在蕴含和矛盾问题上表现尚可,但在中立问题上仍存在系统性转换错误,整体表现有限。研究提示当前模型从图表中获益甚微。

原文摘要 · Abstract (English)

Diagrams are widely used to support logical reasoning, and prior studies suggest that representations such as Euler diagrams can improve human reasoning performance. Recent work has also explored their effects on large language models (LLMs). In this paper, we compare four representational conditions for syllogistic reasoning: natural language, logical notation, linear diagrams, and Euler diagrams. Using 285 problems from Ando et al. (2024), we evaluate two contemporary LLMs, Claude 3.5~Sonnet and GPT-4o-mini. Our results show that diagrammatic representations do not consistently improve performance. Although the models perform well on entailment and contradiction problems, they struggle with neutral problems and often make systematic conversion errors. Overall, the results suggest that the tested models gain limited benefit from diagrams in logical reasoning tasks.

逻辑推理大模型可视化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。