构建分层几何推理评估框架,揭示模型在复杂问题中的真实短板
GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation
- 设计四层递进式推理任务:从视觉感知到自我反思
- 发现子目标分解和无关信息过滤显著提升解题准确率
- 链式思考提示反而降低部分任务表现,适合特定场景
几何问题求解是数学推理的重要分支,依赖对形状与空间关系的精准分析。当前视觉-语言模型的几何推理评估存在局限:测试数据可能受教材内容污染、过度关注最终答案而忽视推理过程、诊断粒度不足。为此,我们提出GeoBench,一个包含四个推理层级的分层基准:视觉感知、目标导向规划、严谨定理应用、自我反思回溯。通过TrustGeoGen生成的六个形式化验证任务,系统评估从属性提取到逻辑错误修正的能力。实验表明,尽管推理模型如OpenAI-o3优于通用多模态大模型,但性能随任务复杂度显著下降。关键发现:子目标分解与无关前提过滤对解题准确率影响重大;而链式思考提示在某些任务中反而降低表现。这些结果确立了GeoBench作为全面评估工具的同时,为几何推理系统的设计提供了可操作指导。
原文摘要 · Abstract (English)
Geometric problem solving constitutes a critical branch of mathematical reasoning, requiring precise analysis of shapes and spatial relationships. Current evaluations of geometric reasoning in vision-language models (VLMs) face limitations, including the risk of test data contamination from textbook-based benchmarks, overemphasis on final answers over reasoning processes, and insufficient diagnostic granularity. To address these issues, we present GeoBench, a hierarchical benchmark featuring four reasoning levels in geometric problem-solving: Visual Perception, Goal-Oriented Planning, Rigorous Theorem Application, and Self-Reflective Backtracking. Through six formally verified tasks generated via TrustGeoGen, we systematically assess capabilities ranging from attribute extraction to logical error correction. Experiments reveal that while reasoning models like OpenAI-o3 outperform general MLLMs, performance declines significantly with increasing task complexity. Key findings demonstrate that sub-goal decomposition and irrelevant premise filtering critically influence final problem-solving accuracy, whereas Chain-of-Thought prompting unexpectedly degrades performance in some tasks. These findings establish GeoBench as a comprehensive benchmark while offering actionable guidelines for developing geometric problem-solving systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。