拆解多模态推理能力,发现模型进步是各子技能此消彼长。
What MLLMs Learn about When they Learn about Multimodal Reasoning
- 用几何题分解感知、推理和多模态特有能力
- 强化学习提升看图能力,文本微调增强反思推理
- 多数错误转向多模态专属问题,揭示评估盲区
多模态推理模型的评估常简化为单一准确率,隐含将推理视为单一能力。我们提出MathLens基准,通过教科书式几何题暴露这一假设:将性能分解为感知、推理和多模态特有三个组件。每道题基于符号规范生成,配有可视化图示、纯文本变体、多模态问题及针对性感知探测,实现对各组件的受控测量。结果显示,常见训练策略导致不同的能力谱系:强化学习主要提升感知基础与对图示变化的鲁棒性,文本SFT则通过反思推理带来提升。随着感知与推理能力增强,剩余错误中越来越多属于多模态特有范畴。这表明,多模态推理的表观进展实为子技能间的平衡转移,而非整体提升,呼吁超越单一准确率的评估方式。
原文摘要 · Abstract (English)
Evaluation of multimodal reasoning models is typically reduced to a single accuracy score, implicitly treating reasoning as a unitary capability. We introduce MathLens, a benchmark of textbook-style geometry problems that exposes this assumption by operationally decomposing performance into perception, reasoning, and multimodal-specific components. Each problem is derived from a symbolic specification and accompanied by visual diagrams, text-only variants, multimodal questions, and targeted perceptual probes, enabling controlled measurement of each component. Using this decomposition, we show that common training strategies induce systematically different capability profiles that are invisible under aggregate accuracy. Reinforcement learning primarily improves perceptual grounding and robustness to diagram variation, while textual SFT yields gains through reflective reasoning. In contrast, as perception and reasoning improve, a growing fraction of remaining errors fall outside these components and are categorized as multimodal-specific. These results suggest that apparent progress in multimodal reasoning reflects shifting balances among subskills rather than uniform advancement, motivating evaluation beyond scalar accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。