TimeMarker让视频理解更准,能精准定位长短视频中的关键时刻。
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
- 用时间分隔符标记视频关键帧,提升时间定位能力
- 支持长短视频动态采样,处理时长跨度大
- 适合需要精确时间定位的视频问答和对话任务
大规模语言模型(LLMs)的发展推动了多模态大语言模型(LMMs)在视觉-语言任务上的进步。然而,现有视频-语言模型常忽视精确的时间定位,且难以处理不同长度的视频。我们提出 TimeMarker,一个面向高质量视频内容对话的通用视频-LLM,强调时间定位能力。TimeMarker 引入时间分隔符标记(Temporal Separator Tokens),增强对视频时间点的感知,准确标记特定时刻。采用 AnyLength 机制实现动态帧采样与自适应标记合并,有效处理短视频与长视频。同时,利用多样数据集,包括重构的时序相关视频问答数据集,强化时间理解能力;结合图像与交错数据,进一步提升语义感知。评估显示,TimeMarker 在多个基准测试中达到领先性能,无论短视频还是长视频均表现优异。
原文摘要 · Abstract (English)
Rapid development of large language models (LLMs) has significantly advanced multimodal large language models (LMMs), particularly in vision-language tasks. However, existing video-language models often overlook precise temporal localization and struggle with videos of varying lengths. We introduce TimeMarker, a versatile Video-LLM designed for high-quality dialogue based on video content, emphasizing temporal localization. TimeMarker integrates Temporal Separator Tokens to enhance temporal awareness, accurately marking specific moments within videos. It employs the AnyLength mechanism for dynamic frame sampling and adaptive token merging, enabling effective handling of both short and long videos. Additionally, TimeMarker utilizes diverse datasets, including further transformed temporal-related video QA datasets, to bolster its temporal understanding capabilities. Image and interleaved data are also employed to further enhance the model's semantic perception ability. Evaluations demonstrate that TimeMarker achieves state-of-the-art performance across multiple benchmarks, excelling in both short and long video categories. Our project page is at \url{https://github.com/TimeMarker-LLM/TimeMarker/}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。