arXiv:2412.05185cs.CVcs.LG2024-12被引 26

将图像大模型轻松升级为视频理解模型,无需重新训练。

LinVT: Empower Your Image-level Large Language Model to Understand Videos

  • 用线性变换保持原模型视觉语言对齐,避免破坏已有能力。
  • 从冗余视频中提炼代表性信息,提升处理效率与准确性。
  • 适配六种主流视觉大模型,性能达当前最佳水平。

大型语言模型(LLMs)已广泛应用于各类任务,促使我们开发基于 LLM 的视频助手。不同于从头训练,我们提出一个模块,可将任意已训练好的图像类 LLM 转化为视频-LLM(经视频数据微调后)。为更好适配图像-LLM 处理视频,我们提出两个设计原则:线性变换以保留原始视觉-语言对齐关系,以及从冗余视频内容中提炼代表性信息。基于此,我们提出即插即用的线性视频分词器(LinVT),使现有图像-LLM 具备视频理解能力。我们在六种近期视觉大模型(Aquila、Blip-3、InternVL2、Mipha、Molmo、Qwen2-VL)上评估 LinVT,验证其高兼容性。基于 LinVT 的 LLM 在多个视频基准测试中达到最先进性能,证明了其在多模态视频理解中的有效性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have been widely used in various tasks, motivating us to develop an LLM-based assistant for videos. Instead of training from scratch, we propose a module to transform arbitrary well-trained image-based LLMs into video-LLMs (after being trained on video data). To better adapt image-LLMs for processing videos, we introduce two design principles: linear transformation to preserve the original visual-language alignment and representative information condensation from redundant video content. Guided by these principles, we propose a plug-and-play Linear Video Tokenizer(LinVT), which enables existing image-LLMs to understand videos. We benchmark LinVT with six recent visual LLMs: Aquila, Blip-3, InternVL2, Mipha, Molmo and Qwen2-VL, showcasing the high compatibility of LinVT. LinVT-based LLMs achieve state-of-the-art performance across various video benchmarks, illustrating the effectiveness of LinVT in multi-modal video understanding.

视频理解大模型迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。