arXiv:2509.24640cs.CVcs.AI2025-09EMNLP被引 2

构建视觉推理新基准SPLICE,揭示大模型在事件理解上的明显短板。

Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs

  • 基于人类筛选的1.1万段视频事件片段,覆盖12类场景
  • 模型准确率远低于人类,尤其在空间与情境推理上差距大
  • 适合研究视觉语言模型推理能力或评估基准设计的学者

本文提出SPLICE,一个源自COIN指令视频数据集的人类筛选基准,用于探测视觉语言模型在时间、因果、空间、上下文及通用知识等多维度上的事件推理能力。SPLICE包含3,381个经人工筛选的视频,覆盖12个类别和180个子类别(如体育、工程、家务),共分割为11,423个事件片段。我们评估了人类与前沿视觉语言模型(VLMs)对这些片段重新排序以形成连贯事件序列的能力。结果显示显著差距:尽管加入人类标注文本可提升模型性能,但不影响人类表现,表明模型更依赖语言先验而非视觉理解。即使有标注,模型仍无法达到人类水平,凸显视觉推理的持续挑战。细分分析显示,模型在时间与因果主导的任务中表现较好,而在空间与情境主导任务中表现较差;日常任务优于专业任务。

原文摘要 · Abstract (English)

In this work, we introduce SPLICE, a human-curated benchmark derived from the COIN instructional video dataset, designed to probe event-based reasoning across multiple dimensions: temporal, causal, spatial, contextual, and general knowledge. SPLICE includes 3,381 human-filtered videos spanning 12 categories and 180 sub-categories, such as sports, engineering, and housework. These videos are segmented into a total of 11,423 event clips. We evaluate both human participants and state-of-the-art vision-language models (VLMs) on the task of rearranging these clips into coherent event sequences to assess visual reasoning capabilities. Results reveal a significant gap: VLMs struggle to match human performance. While human-annotated textual descriptions improve model accuracy, they do not affect human performance, suggesting that models rely more on language priors than on visual understanding. Even with annotations, VLMs fall short of human-level reasoning, underscoring persistent challenges in visual reasoning. A deeper analysis across sub-categories shows that VLMs perform relatively better on videos where temporal and causal reasoning are dominant, compared to those where contextual and spatial reasoning are dominant. They also perform better on everyday tasks than on specialized ones.

视觉推理基准测试VLM事件理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。