arXiv:2505.19853cs.CVcs.AI2025-05NeurIPS被引 1

测试视频模型理解长视频中因果事件的能力,发现现有模型表现不佳。

Two Causally Related Needles in a Video Haystack

  • 设计新基准Causal2Needles,评估模型定位两个事件并理解其因果关系的能力。
  • 模型在相距较远的两个事件上表现更差,性能与距离负相关。
  • 适合研究视频理解、因果推理或长视频建模的学者参考。

评估视频语言模型(VLMs)理解长视频的能力仍具挑战性。我们提出一个名为Causal2Needles的长上下文视频理解基准,用于测试两种现有基准未充分涵盖的关键能力:(1) 从长视频中提取两个独立位置的信息(两个‘针’),并联合理解;(2) 建模人类行为中的因果关系。Causal2Needles通过非因果单针、因果单针和因果双针问题进行评估。最复杂的因果双针问题要求从长视频及对应叙述文本中提取原因和结果事件信息。为避免文本偏见,引入两种互补问题形式:定位包含答案的视频片段,以及描述该片段中的视觉细节。实验表明,当前在已有基准上表现优异的模型,在因果双针问题上仍表现不佳,且模型性能与两个事件间的距离呈负相关。这些发现揭示了当前VLMs的显著局限性。数据集已开源:https://huggingface.co/datasets/causal2needles/Causal2Needles。

原文摘要 · Abstract (English)

Properly evaluating the ability of Video-Language Models (VLMs) to understand long videos remains a challenge. We propose a long-context video understanding benchmark, Causal2Needles, that assesses two crucial abilities insufficiently addressed by existing benchmarks: (1) extracting information from two separate locations (two needles) in a long video and understanding them jointly, and (2) modeling the world in terms of cause and effect in human behaviors. Causal2Needles evaluates these abilities using noncausal one-needle, causal one-needle, and causal two-needle questions. The most complex question type, causal two-needle questions, require extracting information from both the cause and effect events from a long video and the associated narration text. To prevent textual bias, we introduce two complementary question formats: locating the video clip containing the answer, and verbal description of a visual detail from that video clip. Our experiments reveal that models excelling on existing benchmarks struggle with causal 2-needle questions, and the model performance is negatively correlated with the distance between the two needles. These findings highlight critical limitations in current VLMs. The dataset is available at: https://huggingface.co/datasets/causal2needles/Causal2Needles

视频理解因果推理长视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。