为提升模型看图解题能力,提出分阶段处理流程MathFlow
MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems
- 将看图和推理拆成独立阶段,分别优化
- 专用感知模型在多类推理模型上提升显著
- 适用于需要精准理解图表的数学题求解
尽管在多项任务中表现优异,多模态大语言模型(MLLMs)在视觉数学问题求解上仍表现不足,尤其在可靠地感知和解读图表方面。受人类解题思路启发,我们假设从图表中提取有意义信息的能力至关重要,因为它直接决定后续推理。为此,我们提出FlowVerse,一个细粒度评估基准,用于评测MLLMs的感知与推理能力。初步结果表明,现有MLLMs在从图表中提取关键信息和性质,以及基于视觉输入进行复杂推理方面存在明显局限。针对此问题,我们提出MathFlow,一种模块化求解流程,将感知与推理分离,实现独立优化。鉴于当前MLLMs的感知能力有限,我们训练了专门的感知模型MathFlow-P-7B。实验显示,该模型在集成到各类闭源与开源推理模型后均带来显著性能提升,验证了MathFlow流程的有效性及其对不同推理框架的兼容性。
原文摘要 · Abstract (English)
Despite strong results on many tasks, multimodal large language models (MLLMs) still underperform on visual mathematical problem solving, especially in reliably perceiving and interpreting diagrams. Inspired by human problem-solving, we hypothesize that the ability to extract meaningful information from diagrams is pivotal, as it directly conditions subsequent inference. Hence, we introduce FlowVerse, a comprehensive benchmark that provides a fine-grained evaluation of MLLMs' perception and reasoning capabilities. Our preliminary results on FlowVerse reveal that existing MLLMs exhibit substantial limitations when extracting essential information and reasoned properties from diagrams and performing complex reasoning based on these visual inputs. In response, we introduce MathFlow, a modular problem-solving pipeline that decouples perception and inference into distinct stages, thereby optimizing each independently. Given the perceptual limitations observed in current MLLMs, we trained MathFlow-P-7B as a dedicated perception model. Experimental results indicate that MathFlow-P-7B yields substantial performance gains when integrated with various closed-source and open-source inference models. This demonstrates the effectiveness of the MathFlow pipeline and its compatibility with diverse inference frameworks. Project page: https://github.com/MathFlow-zju/MathFlow.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。