arXiv:2509.17437cs.CL2025-09EMNLP被引 11

提升多模态模型几何推理能力,解决视觉感知短板

GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning

  • 分两阶段强化视觉感知再训练推理能力
  • 在几何任务上推理准确率提升9.7%,解题能力提升9.1%
  • 适用于图像理解等视觉密集型任务,适合做多模态模型优化

近期强化学习(RL)进展提升了大语言模型(LLM)的推理能力,但对多模态大模型(MLLM)影响有限。尤其在几何推理等视觉密集型任务中,MLLM常出现幻觉,导致推理错误。我们归因于MLLM的感知瓶颈,限制了推理训练收益。为此,我们设计了面向基础几何概念与空间关系的Geo-Perception问答基准(GeoPQA)。实验表明,MLLM在视觉感知方面存在明显不足,制约了强化学习奖励信号的有效性。为突破该瓶颈,我们提出两阶段强化学习训练框架:先提升几何结构的视觉感知能力,再增强推理能力。应用于Qwen2.5-VL-3B-Instruct模型时,相比直接推理训练,几何推理提升9.7%,问题求解能力提升9.1%。该方法还可推广至图示理解等其他视觉密集型领域,凸显感知基础对有效多模态推理的重要性。

原文摘要 · Abstract (English)

Recent advancements in reinforcement learning (RL) have enhanced the reasoning abilities of large language models (LLMs), yet the impact on multimodal LLMs (MLLMs) is limited. Particularly in vision-intensive tasks like geometric reasoning, MLLMs hallucinate frequently, leading to inaccurate reasoning. We attribute this to the perceptual bottleneck in MLLMs, which caps the benefits of reasoning training. To quantify this, we design a Geo-Perception Question-Answering (GeoPQA) benchmark, targeting basic geometric concepts and spatial relationships. Experiments on GeoPQA reveal significant shortcomings of MLLMs in visual perception, which constrain RL reward signals for effective training. To address this bottleneck, we propose a two-stage RL training framework by first enhancing the visual perception of geometric structures, then fostering reasoning capabilities. Applied to Qwen2.5-VL-3B-Instruct, our two-stage training improves geometric reasoning by 9.7% and geometric problem solving by 9.1%, compared to the direct reasoning training approach. Our method also generalizes to other vision-intensive domains like figure understanding, highlighting the importance of perceptual grounding in effective MLLM reasoning.

多模态模型几何推理强化学习视觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。