arXiv:2603.13091cs.CV2026-03被引 1

评测大模型从视频中抽象推理时空信息的能力,发现现有模型在综合线索上存在明显短板。

Reasoning over Video: Evaluating How MLLMs Extract, Integrate, and Reconstruct Spatiotemporal Evidence

  • 构建新基准VAEX-BENCH,聚焦视频中分散线索的整合与隐含结构推断
  • 在5个抽象推理任务上,主流多模态大模型性能显著低于提取类任务
  • 提供可控制的合成视频数据集,适合研究时空推理与具身智能的学者

随着具身智能的发展,对时空视频理解的需求日益增长,但现有基准主要侧重于可直接提取的推理任务,难以评估多模态大语言模型是否具备抽象时空推理能力。为此,本文提出一个结构化评估框架,系统涵盖抽象时空推理的核心维度,并构建了一个可控的、场景驱动的合成第一视角视频数据集,覆盖物体级、房间级和楼层平面级场景。基于此框架,我们推出了包含五个抽象推理任务及其对应提取任务的基准VAEX-BENCH。通过大量实验对比了先进MLLMs在提取与抽象设置下的表现,揭示其在抽象任务中的局限性,并深入分析了底层瓶颈。数据集即将发布。

原文摘要 · Abstract (English)

The growing interest in embodied agents increases the demand for spatiotemporal video understanding, yet existing benchmarks largely emphasize extractive reasoning, where answers can be explicitly presented within spatiotemporal events. It remains unclear whether multimodal large language models can instead perform abstractive spatiotemporal reasoning, which requires integrating observations over time, combining dispersed cues, and inferring implicit spatial and contextual structure. To address this gap, we formalize abstractive spatiotemporal reasoning from videos by introducing a structured evaluation taxonomy that systematically targets its core dimensions and constructs a controllable, scenario-driven synthetic egocentric video dataset tailored to evaluate abstractive spatiotemporal reasoning capabilities, spanning object-, room-, and floor-plan-level scenarios. Based on this framework, we present VAEX-BENCH, a benchmark comprising five abstractive reasoning tasks together with their extractive counterparts. Our extensive experiments compare the performance of state-of-the-art MLLMs under extractive and abstractive settings, exposing their limitations on abstractive tasks and providing a fine-grained analysis of the underlying bottlenecks. The dataset will be released soon.

视频理解抽象推理多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。