测试大模型对图像序列的时序理解能力,发现差距显著。
Beyond Single Frames: Can LMMs Comprehend Temporal and Contextual Narratives in Image Sequences?
- 构建图像序列任务评估集StripCipher,含三类挑战
- 顶级模型如GPT-4o在重排序任务中仅23.93%准确率
- 揭示大模型在时序推理上的根本性不足,适合研究者参考
大型多模态模型(LMMs)在多种视觉语言任务中取得显著进展,但现有基准主要聚焦单图理解,对图像序列分析仍属空白。为此,我们提出StripCipher,一个全面的基准,用于评估LMMs对图像序列的理解与推理能力。该基准包含人工标注数据集及三个挑战性子任务:视觉叙事理解、上下文帧预测和时序叙事重排序。我们对16个前沿LMMs(包括GPT-4o和Qwen2.5VL)进行评估,结果显示其性能与人类存在显著差距,尤其在打乱图像序列的重排序任务中表现欠佳。例如,GPT-4o在重排序任务中准确率仅为23.93%,比人类低56.07%。进一步定量分析表明,图像输入格式等因素影响模型在序列理解中的表现,凸显了当前LMMs在时序建模方面仍面临根本性挑战。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) have achieved remarkable success across various visual-language tasks. However, existing benchmarks predominantly focus on single-image understanding, leaving the analysis of image sequences largely unexplored. To address this limitation, we introduce StripCipher, a comprehensive benchmark designed to evaluate capabilities of LMMs to comprehend and reason over sequential images. StripCipher comprises a human-annotated dataset and three challenging subtasks: visual narrative comprehension, contextual frame prediction, and temporal narrative reordering. Our evaluation of 16 state-of-the-art LMMs, including GPT-4o and Qwen2.5VL, reveals a significant performance gap compared to human capabilities, particularly in tasks that require reordering shuffled sequential images. For instance, GPT-4o achieves only 23.93% accuracy in the reordering subtask, which is 56.07% lower than human performance. Further quantitative analysis discuss several factors, such as input format of images, affecting the performance of LLMs in sequential understanding, underscoring the fundamental challenges that remain in the development of LMMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。