arXiv:2507.08306cs.AIcs.CL2025-07被引 21

让多模态大模型同时具备通用与空间推理能力

M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning

  • 构建高质量数据集,含29.42万条逻辑连贯的推理样本
  • 在8个基准上达到新SOTA,空间推理能力显著提升
  • 适合需要复杂环境理解的智能体、机器人应用

近期基于可验证奖励强化学习(RLVR)的多模态大语言模型(MLLM)在推理能力上取得显著进展,但仍面临动态空间交互能力不足的问题。为此,我们提出M2-Reasoning-7B模型,兼顾通用与空间推理。方法包括:(1) 构建全新数据流水线,生成294.2万条高质量样本(其中168万用于冷启动微调,126.2万用于RLVR),包含逻辑连贯的推理轨迹并经全面评估;(2) 采用动态多任务训练策略,分步优化以缓解任务冲突,并设计任务专属奖励信号。该组合使模型在8个基准上达到新SOTA,显著提升通用与空间推理表现。

原文摘要 · Abstract (English)

Recent advancements in Multimodal Large Language Models (MLLMs), particularly through Reinforcement Learning with Verifiable Rewards (RLVR), have significantly enhanced their reasoning abilities. However, a critical gap persists: these models struggle with dynamic spatial interactions, a capability essential for real-world applications. To bridge this gap, we introduce M2-Reasoning-7B, a model designed to excel in both general and spatial reasoning. Our approach integrates two key innovations: (1) a novel data pipeline that generates 294.2K high-quality data samples (168K for cold-start fine-tuning and 126.2K for RLVR), which feature logically coherent reasoning trajectories and have undergone comprehensive assessment; and (2) a dynamic multi-task training strategy with step-wise optimization to mitigate conflicts between data, and task-specific rewards for delivering tailored incentive signals. This combination of curated data and advanced training allows M2-Reasoning-7B to set a new state-of-the-art (SOTA) across 8 benchmarks, showcasing superior performance in both general and spatial reasoning domains.

多模态推理空间理解强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。