用生成模型主动构建视频评测数据,精准测试多模态大模型的时空推理能力。
VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis

- 通过多智能体管道生成可控多样视频,实现主动合成评测场景。
- 建立3×2×2视频分类体系,覆盖空间尺度、视角与场景动态多样性。
- 设计分层任务,分离视觉感知与高层时空推理,便于精准诊断模型短板。
时空推理是多模态大语言模型在真实世界中运行的核心能力,其精确评估已成为关键挑战。现有基准数据集主要依赖静态图像集或被动收集的视频数据,难以评估细粒度的推理能力。本文提出VGenST-Bench,一个基于生成模型主动合成高度可控且多样评测场景的视频基准。为构建该基准,我们设计了包含人工质量控制环节的多智能体流水线,确保所有生成视频与问答对的质量。我们建立了全面的3×2×2视频分类体系,涵盖空间尺度、视角与场景动态,以覆盖多样化场景。此外,我们设计了分层任务套件,将低层级视觉感知与高层时空推理解耦。通过从被动整理转向主动合成,VGenST-Bench实现了对多模态大模型时空理解能力的精细化诊断。
原文摘要 · Abstract (English)
Spatio-temporal reasoning is a core capability for Multimodal Large Language Models (MLLMs) operating in the real world. As such, evaluating it precisely has become an essential challenge. However, existing spatio-temporal reasoning benchmark datasets primarily rely on static image sets or passively curated video data, which limits the evaluation of fine-grained reasoning capabilities. In this paper, we introduce VGenST-Bench, a video benchmark that employs generative models to actively synthesize highly controlled and diverse evaluation scenarios. To construct VGenST-Bench, we propose a multi-agent pipeline incorporating a human quality control stage, ensuring the quality of all generated videos and QA pairs. We establish a comprehensive 3x2x2 video taxonomy, encompassing Spatial Scale, Perspective, and Scene Dynamics to span diverse scenarios. Furthermore, we design a hierarchical task suite that decouples low-level visual perception from high-level spatio-temporal reasoning. By shifting the paradigm from passive curation to active synthesis, VGenST-Bench enables fine-grained diagnosis of spatio-temporal understanding in MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。