构建首个细粒度时间对齐的长视频摘要基准,评估模型时间感知能力。
LVSum: A Benchmark for Timestamp-Aware Long Video Summarization

- 设计人类标注的多领域长视频摘要数据集,支持时间锚定。
- 发现语音转录比视觉帧对摘要质量贡献更大,模型仍落后于人工。
- 揭示当前多模态大模型在时间定位、指令遵循与跨模态一致上的系统性缺陷。
长视频摘要对多模态大语言模型(MLLMs)构成重大挑战,尤其在长时间跨度下保持时间准确性及语义-时间双重锚定。本文提出LVSum,一个基于人工标注的细粒度时间对齐长视频摘要评测基准。该数据集包含72个涵盖13个领域的多样化视频,平均时长16分钟,每段视频配有最多10份带时间标记的人类摘要。我们采用新提出的基于LLM的指标评估内容相关性与模态一致性,并结合标准自动指标对主流开源与闭源MLLM进行综合评测。实验揭示三个关键发现:(1) 语音转录比仅依赖视觉帧显著提升摘要质量;(2) 模型生成摘要与人工摘要间存在显著性能差距;(3) 当前MLLM在时间定位、指令遵从和跨模态一致性方面存在系统性弱点。数据集与代码已公开。
原文摘要 · Abstract (English)
Long video summarization presents significant challenges for multimodal large language models (MLLMs), particularly in maintaining temporal fidelity over extended durations and producing summaries that are both semantically and temporally grounded. We introduce LVSum, a human-annotated benchmark for evaluating long-form video summarization with fine-grained temporal alignment. LVSum comprises 72 diverse videos spanning 13 domains with an average duration of 16 minutes, each annotated with up to 10 human-generated summaries containing temporal references. We conduct a comprehensive evaluation of leading proprietary and open-source MLLMs using newly introduced LLM-based metrics for content relevance and modality coherence, alongside standard automatic metrics. Our experiments reveal three key findings: (1) transcripts contribute substantially more to summarization quality than visual frames alone, (2) a significant performance gap persists between model-generated and human-written summaries, and (3) current MLLMs exhibit systematic weaknesses in temporal grounding, instruction adherence, and cross-modal coherence. We release the dataset and code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。