提升数学视觉理解能力,让大模型更准识别几何图形。
Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs
- 设计新视觉编码器与特征路由机制,精准捕捉几何细节
- 在数学数据集上比7B模型高15%准确率,错误率降低至30%以下
- 适合需要精细视觉推理的数学智能研究者使用
当前多模态大模型在依赖细粒度视觉理解的数学任务中表现不佳,主要源于图像级对比预训练(如CLIP)对几何基本元素感知不足。本文系统评估了先进多模态大模型的视觉定位能力,发现其视觉定位准确率与解题性能存在显著负相关,突显细粒度视觉理解的关键作用。值得注意的是,即使先进模型GPT-4o在识别几何实体时仍存在70%错误率。为此,提出SVE-Math(选择性视觉增强型数学多模态模型),采用基于几何的视觉编码器和动态特征路由机制,能准确识别视觉基本元素,并生成适配语言模型推理需求的提示。实验表明,SVE-Math-Qwen2.5-7B在MathVerse上优于其他7B模型15%,且兼容GPT-4V在MathVista上的表现。尽管训练数据量较小,其在GeoQA上表现媲美在更大数据集上训练的模型。研究强调将细粒度视觉理解融入多模态模型的重要性,为未来研究提供新方向。
原文摘要 · Abstract (English)
Current multimodal large language models (MLLMs) often underperform on mathematical problem-solving tasks that require fine-grained visual understanding. The limitation is largely attributable to inadequate perception of geometric primitives during image-level contrastive pre-training (e.g., CLIP). While recent efforts to improve math MLLMs have focused on scaling up mathematical visual instruction datasets and employing stronger LLM backbones, they often overlook persistent errors in visual recognition. In this paper, we systematically evaluate the visual grounding capabilities of state-of-the-art MLLMs and reveal a significant negative correlation between visual grounding accuracy and problem-solving performance, underscoring the critical role of fine-grained visual understanding. Notably, advanced models like GPT-4o exhibit a 70% error rate when identifying geometric entities, highlighting that this remains a key bottleneck in visual mathematical reasoning. To address this, we propose a novel approach, SVE-Math (Selective Vision-Enhanced Mathematical MLLM), featuring a geometric-grounded vision encoder and a feature router that dynamically adjusts the contribution of hierarchical visual feature maps. Our model recognizes accurate visual primitives and generates precise visual prompts tailored to the language model's reasoning needs. In experiments, SVE-Math-Qwen2.5-7B outperforms other 7B models by 15% on MathVerse and is compatible with GPT-4V on MathVista. Despite being trained on smaller datasets, SVE-Math-7B achieves competitive performance on GeoQA, rivaling models trained on significantly larger datasets. Our findings emphasize the importance of incorporating fine-grained visual understanding into MLLMs and provide a promising direction for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。