评测大模型在教学视频中的跨模态时序理解能力,发现顶尖模型仍有明显短板。
LEMON: How Well Do MLLMs Perform Temporal Multimodal Understanding on Instructional Videos?
- 构建覆盖5个学科29门课的长视频评测基准LEMON,含4181个高质量问答对。
- 模型在时序推理和教学内容预测任务上表现不佳,即使GPT-4o也存在明显差距。
- 适合研究长视频理解、多模态推理与教育AI的学者使用。
近期多模态大模型在视觉、音频和语言任务上取得显著进展,但其在长时序、知识密集型且具有时序结构的教育类内容上的表现仍缺乏研究。为此,我们提出LEMON——基于讲座的多模态理解评测基准,聚焦需长时程推理与跨模态整合的STEM教学视频。LEMON包含2,277段视频片段,覆盖5个学科、29门课程,平均时长196.1秒,共生成4,181个高质量问答对(3,413个选择题,768个开放题)。该基准具有:(1) 语义丰富度与学科密度高,(2) 视频-音频-文本模态紧密耦合,(3) 明确的时间与教学结构,(4) 上下文关联的多轮提问。涵盖六大核心任务与十二个子任务,覆盖从感知到推理再到生成的全认知链条。全面实验揭示各任务间存在显著性能差距,表明即使是最先进的模型如GPT-4o,在时序推理与教学预测方面仍表现不足。我们期望LEMON成为可扩展、具挑战性的基准,推动长时教育内容中多模态感知、推理与生成的发展。
原文摘要 · Abstract (English)
Recent multimodal large language models (MLLMs) have shown remarkable progress across vision, audio, and language tasks, yet their performance on long-form, knowledge-intensive, and temporally structured educational content remains largely unexplored. To bridge this gap, we introduce LEMON, a Lecture-based Evaluation benchmark for MultimOdal uNderstanding, focusing on STEM lecture videos that require long-horizon reasoning and cross-modal integration. LEMON comprises 2,277 video segments spanning 5 disciplines and 29 courses, with an average duration of 196.1 seconds, yielding 4,181 high-quality QA pairs, including 3,413 multiple-choice and 768 open-ended questions. Distinct from existing video benchmarks, LEMON features: (1) semantic richness and disciplinary density, (2) tightly coupled video-audio-text modalities, (3) explicit temporal and pedagogical structure, and (4) contextually linked multi-turn questioning. It further encompasses six major tasks and twelve subtasks, covering the full cognitive spectrum from perception to reasoning and then to generation. Comprehensive experiments reveal substantial performance gaps across tasks, highlighting that even state-of-the-art MLLMs like GPT-4o struggle with temporal reasoning and instructional prediction. We expect LEMON to serve as an extensible and challenging benchmark for advancing multimodal perception, reasoning, and generation in long-form instructional contents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。