arXiv:2506.10415cs.CLcs.CV2025-06ACL被引 8

测试大模型理解图像序列时间顺序的能力,发现普遍表现不佳。

Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?

  • 构建多模态大模型时间推理基准TempVS,包含事件关系、句子与图像排序任务。
  • 38个主流模型在任务上表现远低于人类,最高准确率不足40%。
  • 适合研究视觉语言模型时序理解的学者,推动模型改进方向。

本文提出TempVS基准,聚焦多模态大语言模型(MLLMs)在图像序列中的时间定位与推理能力。TempVS包含三项核心测试(事件关系推断、句子排序、图像排序),每项均配有基础定位测试。该任务要求模型结合视觉与语言模态,理解事件的时间顺序。我们评估了38个前沿MLLMs,结果显示模型在该任务上表现不佳,与人类能力存在显著差距。研究还提供了细粒度分析,揭示未来研究的潜在方向。TempVS的数据与代码已开源:https://github.com/yjsong22/TempVS。

原文摘要 · Abstract (English)

This paper introduces the TempVS benchmark, which focuses on temporal grounding and reasoning capabilities of Multimodal Large Language Models (MLLMs) in image sequences. TempVS consists of three main tests (i.e., event relation inference, sentence ordering and image ordering), each accompanied with a basic grounding test. TempVS requires MLLMs to rely on both visual and linguistic modalities to understand the temporal order of events. We evaluate 38 state-of-the-art MLLMs, demonstrating that models struggle to solve TempVS, with a substantial performance gap compared to human capabilities. We also provide fine-grained insights that suggest promising directions for future research. Our TempVS benchmark data and code are available at https://github.com/yjsong22/TempVS.

多模态时间推理基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。