arXiv:2409.12963cs.CVcs.AI2024-09被引 13

无需训练即可让视频大模型处理更长视频,解决扩展性难题。

Interpolating Video-LLMs: Toward Longer-sequence LMMs in a Training-free Manner

  • 通过重排视频帧令牌,绕过固定编码器限制。
  • 不需微调即可扩展语言模型上下文窗口,支持更多视频帧。
  • 适合希望低成本扩展视频理解能力的研究者与开发者。

大型语言模型(LLMs)的发展推动了多种视频模态融合策略。其中,视频-语言模型(Video-LLMs)通过可优化的接口将复杂视频编码器与语言模型连接。然而,受计算和数据限制,这些模型通常仅预训练用于处理短视频,难以应用于长视频内容理解。此外,微调以支持更长视频成本过高。因此,探索完全无需训练的视频-语言模型插值方法至关重要。本文首先识别两大挑战:(1) 视频编码器与模态对齐投影器固定,无法整合额外帧;(2) 语言模型主干在内容长度上受限,难以处理增加的视频令牌。为此,我们提出一种名为INTP-Video-LLMs的特定插值方法。引入替代的视频令牌重排技术,规避固定编码器与投影器的限制;同时,提出一种无需训练的语言模型上下文窗口扩展方法,使模型能理解相应增加的视觉令牌数量。

原文摘要 · Abstract (English)

Advancements in Large Language Models (LLMs) inspire various strategies for integrating video modalities. A key approach is Video-LLMs, which incorporate an optimizable interface linking sophisticated video encoders to LLMs. However, due to computation and data limitations, these Video-LLMs are typically pre-trained to process only short videos, limiting their broader application for understanding longer video content. Additionally, fine-tuning Video-LLMs to handle longer videos is cost-prohibitive. Consequently, it becomes essential to explore the interpolation of Video-LLMs under a completely training-free setting. In this paper, we first identify the primary challenges in interpolating Video-LLMs: (1) the video encoder and modality alignment projector are fixed, preventing the integration of additional frames into Video-LLMs, and (2) the LLM backbone is limited in its content length capabilities, which complicates the processing of an increased number of video tokens. To address these challenges, we propose a specific INTerPolation method for Video-LLMs (INTP-Video-LLMs). We introduce an alternative video token rearrangement technique that circumvents limitations imposed by the fixed video encoder and alignment projector. Furthermore, we introduce a training-free LLM context window extension method to enable Video-LLMs to understand a correspondingly increased number of visual tokens.

视频生成大模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。