arXiv:2603.00515cs.CV2026-03被引 7

让大模型从2D视频看懂3D空间随时间变化,突破视觉时空推理瓶颈。

MLLM-4D: Towards Visual-based Spatial-Temporal Intelligence

  • 用立体视频数据生成高质量4D指令数据,构建新训练集
  • 仅通过后训练策略实现顶尖时空理解与推理能力
  • 适合研究多模态智能、视频理解与具身认知的学者

人类天生具备基于视觉的四维时空智能,能仅凭视觉输入感知和推理三维空间随时间的演变。尽管至关重要,这一能力仍是当前多模态大模型(MLLMs)的重大瓶颈。为此,我们提出MLLM-4D框架,旨在解决训练数据构建与模型后训练中的关键缺口。在数据层面,设计低成本的数据构建流程,将现有立体视频数据集转化为高质量的4D时空指令数据,形成用于监督微调(SFT)的MLLM4D-2M和用于强化学习微调(RFT)的MLLM4D-R1-30k数据集,以及用于全面评估的MLLM4D-Bench。在模型训练方面,通过SFT建立基础4D理解,并利用组相对策略优化(GRPO)结合专有的时空思维链(ST-CoT)提示与时空奖励函数(ST-reward),无需修改模型架构即显著提升4D推理能力。大量实验表明,MLLM-4D仅依赖2D RGB输入即可实现当前最优的时空理解与推理性能。

原文摘要 · Abstract (English)

Humans are born with vision-based 4D spatial-temporal intelligence, which enables us to perceive and reason about the evolution of 3D space over time from purely visual inputs. Despite its importance, this capability remains a significant bottleneck for current multimodal large language models (MLLMs). To tackle this challenge, we introduce MLLM-4D, a comprehensive framework designed to bridge the gaps in training data curation and model post-training for spatiotemporal understanding and reasoning. On the data front, we develop a cost-efficient data curation pipeline that repurposes existing stereo video datasets into high-quality 4D spatiotemporal instructional data. This results in the MLLM4D-2M and MLLM4D-R1-30k datasets for Supervised Fine-Tuning (SFT) and Reinforcement Fine-Tuning (RFT), alongside MLLM4D-Bench for comprehensive evaluation. Regarding model training, our post-training strategy establishes a foundational 4D understanding via SFT and further catalyzes 4D reasoning capabilities by employing Group Relative Policy Optimization (GRPO) with specialized Spatiotemporal Chain of Thought (ST-CoT) prompting and Spatiotemporal reward functions (ST-reward) without involving the modification of architecture. Extensive experiments demonstrate that MLLM-4D achieves state-of-the-art spatial-temporal understanding and reasoning capabilities from purely 2D RGB inputs. Project page: https://github.com/GVCLab/MLLM-4D.

多模态时空推理视频理解大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。