arXiv:2510.10976cs.AI2025-10被引 5

用图结构强化视频时空推理,让大模型更懂物体位置与运动关系。

Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph

  • 构建图结构引导模型推理场景的时空拓扑关系
  • 在STI-Bench上比基础模型提升13%准确率
  • 适合需要精确时空理解的机器人、VR等应用

多模态大语言模型在语义理解上表现良好,但在精确的时空理解方面仍存在不足。现有方法多关注视频内容本身,忽视了视频中的物理信息,如多物体布局和运动规律。为解决此问题,我们提出Video-STR,一种基于图结构的强化学习方法,用于提升视频时空推理能力。该方法利用可验证奖励的强化学习(RLVR)框架,引入基于图的组相对策略优化(GRPO)机制,在思考过程中引导模型推断场景的底层时空结构。为弥补时空训练数据不足,我们构建了包含20.5万组问答对的STV-205k数据集,涵盖室内外动态多物体场景。实验表明,Video-STR在多个基准测试中达到领先水平,相较于基线模型在STI-Bench上提升13%。代码、模型与数据将公开发布。

原文摘要 · Abstract (English)

Recent progress in Multimodal Large Language Models (MLLMs) has demonstrated strong semantic understanding capabilities, but struggles to perform precise spatio-temporal understanding. Existing spatio-temporal methods primarily focus on the video itself, while overlooking the physical information within the video, such as multi-object layouts and motion. Such limitations restrict the use of MLLMs in downstream applications that demand high precision, including embodied intelligence and VR. To address this issue, we present Video-STR, a novel graph-based reinforcement method for precise Video Spatio-Temporal Reasoning. Building upon the capacity of Reinforcement Learning with Verifiable Reward (RLVR) to improve model abilities, we introduce a reasoning mechanism using graph-based Group Relative Policy Optimization (GRPO) method to guide the model in inferring the underlying spatio-temporal topology of scenarios during the thinking process. To resolve the lack of spatio-temporal training data, we construct the STV-205k dataset with 205k question-answering pairs, covering dynamic multi-object scenes in both indoor and outdoor environments, to support the model training. Experiments show that Video-STR achieves state-of-the-art results on various benchmarks, outperforming the base model by 13% on STI-Bench, and demonstrating the effectiveness of our approach and dataset. Code, model, and data will be released.

视频推理图神经网络多模态模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。