arXiv:2510.09302cs.CVcs.AI2025-10被引 5

用文字描述图来提升大模型解几何题能力

CapGeo: A Caption-Assisted Approach to Geometric Reasoning

  • 把几何图转成精准文字描述,辅助模型推理
  • 某模型准确率从8.6%升至59.0%,另一从44.8%升至73.0%
  • 新评测集支持精准评估图转文质量,适合视觉推理研究者

几何推理仍是多模态大模型的核心挑战。即使最先进的闭源系统如GPT-O3和Gemini-2.5-Pro,在国际数学奥林匹克竞赛等任务中表现出强文本推理能力,仍难以可靠解决几何问题。这表明瓶颈在于理解几何图示,而非推理本身。由于几何图形常能以简洁文本忠实描述,将视觉内容转化为文字说明是一条可行路径。为此,我们提出CapGeo——一种基于图文描述的推理框架,实现视觉与文本模态的桥梁。实验显示,引入文字描述后性能显著提升:Qwen2.5-VL-72B从8.6%(仅视觉)提升至59.0%,Claude-Opus-4从44.8%升至73.0%。为进一步系统评估并筛选高质量几何图转文模型,我们构建了包含4,641对精选图-文配对的CapGeo-Bench数据集。该数据集采用关键点匹配评估指标,与下游CapGeo性能高度相关,可有效衡量几何图转文能力。整体框架与评测集为推进多模态大模型几何推理提供了新路径。

原文摘要 · Abstract (English)

Geometric reasoning remains a core challenge for Multimodal Large Language Models (MLLMs). Even the most advanced closed-source systems, such as GPT-O3 and Gemini-2.5-Pro, still struggle to solve geometry problems reliably, despite exhibiting strong textual reasoning abilities on tasks like the International Mathematical Olympiad (IMO). This gap suggests that the bottleneck lies in understanding geometric diagrams rather than reasoning itself. Since geometric figures can often be faithfully described in concise textual form, converting visual content into captions offers a promising direction. Motivated by this insight, we introduce CapGeo, a caption-assisted reasoning framework that bridges visual and textual modalities. Experiments show substantial improvements when models are equipped with captions: Qwen2.5-VL-72B improves from 8.6% (vision-only) to 59.0%, while Claude-Opus-4 rises from 44.8% to 73.0%. To systematically evaluate and identify high-quality geometric captioning models, we further propose CapGeo-Bench, a dataset of 4,641 curated figure-caption pairs. Crucially, CapGeo-Bench incorporates a keypoint-based evaluation metric that correlates strongly with downstream CapGeo performance, enabling reliable assessment of geometric captioning ability. Together, our framework and benchmark highlight a new pathway toward advancing geometric reasoning in MLLMs.

几何推理多模态图文生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。