构建9万道带图几何题,评估大模型的多步推理能力
GeoChallenge: A Multi-Answer Multiple-Choice Benchmark for Geometric Reasoning with Diagrams
- 自动生成9万道需图文结合的多步几何证明题
- 人类准确率94.74%,顶尖模型仅75.89%仍有差距
- 揭示模型三大缺陷:选错、不看图、乱推演
评估大语言模型(LLMs)的符号推理能力需要依赖包含文本与图表的几何基准。然而,现有基准规模有限,且很少提供视觉引导的多选题,难以可靠评估复杂推理能力。我们提出GeoChallenge,一个包含9万道自动生成的多选几何证明题的数据集,每道题均需基于对齐的文本描述与图表进行多步推理。该数据集提供细粒度复杂度评级和形式化语言标注,支持可控评估。在多个先进大模型上的实验表明,模型与人类之间存在明显差距:表现最好的模型GPT-5-nano在精确匹配上达到75.89%,而人类为94.74%。进一步分析揭示了模型的三种常见失败模式:(1)在多选设置下出现精确匹配错误;(2)对视觉信息依赖弱;(3)推理过度发散但无法收敛。
原文摘要 · Abstract (English)
Evaluating the symbolic reasoning of large language models (LLMs) calls for geometry benchmarks that require multi-step proofs grounded in both text and diagrams. However, existing benchmarks are often limited in scale and rarely provide visually grounded multiple-choice questions, limiting reliable evaluation of complex reasoning. We introduce GeoChallenge, a dataset of 90K automatically generated multiple-choice geometry proof problems, each requiring multi-step reasoning over aligned textual descriptions and diagrams. GeoChallenge provides fine-grained complexity ratings and formal language annotations to enable controlled evaluation. Experiments on multiple advanced LLMs show a clear performance gap between models and humans (the best-performing model, GPT-5-nano, achieves 75.89 exact match vs. 94.74 for humans). Further analysis also reveals three common failure patterns of LLMs: (1) exact match failures under the multiple-choice setting; (2) weak visual reliance; and (3) overextended reasoning without convergence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。