arXiv:2410.04511cs.CVcs.CL2024-10被引 2

不需微调,用多模型协作实现更懂语义的视频摘要

Realizing Video Summarization from the Path of Language-based Semantic Understanding

  • 引入专家混合机制,融合多个VideoLLM优势
  • 无需微调即可生成连贯详实的文本摘要
  • 适合需要语义理解与灵活扩展的视频摘要场景

近年来,基于视频的大语言模型(VideoLLMs)通过将视频特征乃至部分音频特征与大语言模型对齐,显著推动了视频摘要的发展。然而,每种VideoLLM均有其优劣。许多现有方法需大量微调以弥补缺陷,资源消耗高。本文观察到不同VideoLLM的优势可互补,提出一种受专家混合(MoE)范式启发的新型视频摘要框架,该框架为推理时算法,无需任何微调。通过集成多个VideoLLMs,生成全面且连贯的文本摘要,有效融合视觉与音频内容,提供详细背景描述,并在关键帧识别上表现优异,相比仅依赖视觉的传统计算机视觉方法,实现更具语义意义的检索。此外,生成的摘要还能提升下游任务如摘要视频生成的性能,可通过关键帧选择或结合文生图模型实现。本语言驱动方法提供了传统方法的语义丰富替代方案,具备灵活融入新VideoLLMs的能力,增强了视频摘要任务的适应性与性能。

原文摘要 · Abstract (English)

The recent development of Video-based Large Language Models (VideoLLMs), has significantly advanced video summarization by aligning video features and, in some cases, audio features with Large Language Models (LLMs). Each of these VideoLLMs possesses unique strengths and weaknesses. Many recent methods have required extensive fine-tuning to overcome the limitations of these models, which can be resource-intensive. In this work, we observe that the strengths of one VideoLLM can complement the weaknesses of another. Leveraging this insight, we propose a novel video summarization framework inspired by the Mixture of Experts (MoE) paradigm, which operates as an inference-time algorithm without requiring any form of fine-tuning. Our approach integrates multiple VideoLLMs to generate comprehensive and coherent textual summaries. It effectively combines visual and audio content, provides detailed background descriptions, and excels at identifying keyframes, which enables more semantically meaningful retrieval compared to traditional computer vision approaches that rely solely on visual information, all without the need for additional fine-tuning. Moreover, the resulting summaries enhance performance in downstream tasks such as summary video generation, either through keyframe selection or in combination with text-to-image models. Our language-driven approach offers a semantically rich alternative to conventional methods and provides flexibility to incorporate newer VideoLLMs, enhancing adaptability and performance in video summarization tasks.

视频摘要多模型融合大模型应用语义理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。