arXiv:2605.01391cs.CV2026-05中稿 · CVPR

首个面向视频交互的时空理解评测基准,专为复杂多实体互动设计。

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark

论文配图:VISTA: Video Interaction Spatio-Temporal Analysis Benchmark
图 1 · 摘自论文原文
  • 将视频分解为实体、动作与关系动态,支持多轴诊断
  • 覆盖1.2万条视频-查询对,涵盖多样场景与复杂度
  • 揭示主流模型在时空理解上的隐藏偏差,指导模型优化

现有视觉语言模型(VLM)评测基准多聚焦于简单单动作视频、封闭属性集和有限实体类型,难以反映真实世界视频中自由形式、多动作、多实体交互的复杂性。同时,缺乏系统性框架来分析模型在时空不同维度上的失败模式。为此,我们提出VISTA——一个面向开放集、多实体、多动作的视频交互时空理解评测基准。VISTA将视频分解为可解释的实体、其关联动作及关系动态,支持对关系、空间与时间理解的多轴诊断与统一评估。基准整合多个数据集,构建统一的交互感知分类体系,包含约1.2万条精心筛选的视频-查询对,覆盖多样化场景与复杂程度。我们系统评估了11个先进VLMs,并基于分类体系拆解性能,揭示传统指标掩盖的缺陷与显著的时空偏见。VISTA提供了在挑战性数据集上的细粒度、分类驱动诊断,为模型设计、预训练策略与评估协议提供精细指引。总体而言,VISTA是首个大规模、交互感知的时空理解诊断基准。

原文摘要 · Abstract (English)

Existing benchmarks for Vision-Language Models (VLMs) primarily evaluate spatio-temporal understanding on simple single-action videos, closed attribute sets and restricted entity types, failing to capture the freeform, multi-action interactions between diverse entities which characterize real-world video understanding. Furthermore, the lack of a systematic framework for analyzing model failures across complementary spatio-temporal axes hinders comprehensive evaluation. To address these gaps, we introduce VISTA, a Video Interaction Spatio-Temporal Analysis benchmark designed for open-set, multi-entity and multi-action spatio-temporal understanding in VLMs. VISTA decomposes videos into interpretable entities, their associated actions, and relational dynamics, enabling multi-axis diagnostics and unified assessment of relational, spatial, and temporal understanding. Our benchmark integrates multiple datasets into a single interaction-aware taxonomy and comprises ~12K curated video-query pairs spanning diverse scenes and complexities. We systematically evaluate 11 state-of-the-art VLMs on VISTA, and break down aggregate performance across our taxonomy to reveal shortcomings and pronounced spatio-temporal biases obscured by traditional metrics. By providing detailed, taxonomy-driven diagnostics on a challenging dataset, VISTA offers a nuanced framework to guide advances in model design, pretraining strategies, and evaluation protocols. Overall, VISTA is the first, large-scale, interaction-aware diagnostic benchmark for spatio-temporal understanding in VLMs.

视频理解多模态评测基准时空分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。