让多模态大模型像人一样理解第一视角视频中的时空动态。
ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos
- 用逆向思维增强强化学习,提升模型推理能力。
- 在5000+问答数据集上显著优于基线模型。
- 适合研究视频理解与具身智能的学者参考。
人类能轻松从第一视角视频中理解动态视觉事件,但多模态大语言模型(MLLMs)是否具备类似时空推理能力尚不明确。本文探索第一视角下的多模态时空推理,旨在赋予MLLMs类人推理能力。为此,我们构建了包含超过5000个问答对的Ego-ST Bench基准,涵盖空间、时间及融合时空推理四类任务,系统评估模型表现。同时提出ST-R1训练范式,将逆向思维引入强化学习过程,结合长链式思维(long-CoT)监督微调与组相对策略优化(GRPO),仅用有限高质量数据即实现显著性能提升。Ego-ST Bench与ST-R1为视频驱动的时空推理研究提供了重要资源与洞见。
原文摘要 · Abstract (English)
Humans excel at spatial-temporal reasoning, effortlessly interpreting dynamic visual events from an egocentric viewpoint. However, whether multimodal large language models (MLLMs) can similarly understand the 4D world remains uncertain. This paper explores multimodal spatial-temporal reasoning from an egocentric perspective, aiming to equip MLLMs with human-like reasoning capabilities. To support this objective, we introduce \textbf{Ego-ST Bench}, a novel benchmark containing over 5,000 question-answer pairs across four categories, systematically evaluating spatial, temporal, and integrated spatial-temporal reasoning. Additionally, we propose \textbf{ST-R1} training paradigm, a video-based reasoning model that incorporates reverse thinking into its reinforcement learning process, significantly enhancing performance. We combine long-chain-of-thought (long-CoT) supervised fine-tuning with Group Relative Policy Optimization (GRPO) reinforcement learning, achieving notable improvements with limited high-quality data. Ego-ST Bench and ST-R1 provide valuable insights and resources for advancing video-based spatial-temporal reasoning research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。