图像训练的视频大模型也能很好进行时间推理,视频数据提升有限。
How Important are Videos for Training Video LLMs?
- 用图像序列+问答微调,模拟视频能力
- 图像训练模型在时间推理任务上显著优于随机猜测
- 当前视频训练效率低,存在关键瓶颈待研究
视频大语言模型(Video LLM)研究快速发展,多数模型基于预训练文本模型,通过图像和视频字幕数据微调。本文发现,仅用图像训练的模型在时间推理任务上表现远超预期,视频微调带来的性能提升非常有限。我们验证了两种基于LongVU算法的模型,在仅图像训练下于TVBench基准上显著高于随机水平。此外,提出一种简单的微调方案:使用带标注的图像序列与时间相关问题,其性能接近甚至超过视频训练模型。这表明当前模型对真实视频中的丰富时序特征利用不足。研究呼吁深入探究图像训练模型实现时间推理的机制,以及现有视频训练方案效率低下的根本原因。
原文摘要 · Abstract (English)
Research into Video Large Language Models (LLMs) has progressed rapidly, with numerous models and benchmarks emerging in just a few years. Typically, these models are initialized with a pretrained text-only LLM and finetuned on both image- and video-caption datasets. In this paper, we present findings indicating that Video LLMs are more capable of temporal reasoning after image-only training than one would assume, and that improvements from video-specific training are surprisingly small. Specifically, we show that image-trained versions of two LLMs trained with the recent LongVU algorithm perform significantly above chance level on TVBench, a temporal reasoning benchmark. Additionally, we introduce a simple finetuning scheme involving sequences of annotated images and questions targeting temporal capabilities. This baseline results in temporal reasoning performance close to, and occasionally higher than, what is achieved by video-trained LLMs. This suggests suboptimal utilization of rich temporal features found in real video by current models. Our analysis motivates further research into the mechanisms that allow image-trained LLMs to perform temporal reasoning, as well as into the bottlenecks that render current video training schemes inefficient.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。