用文本引导筛选视频时空信息,高效理解长视频内容
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering
- 根据用户提问动态过滤无关视觉片段,仅保留关键帧
- 仅用16个视觉标记即达零样本领先效果,训练数据少10倍
- 适合需要精准定位长视频内容的场景,如智能客服、视频检索
近期多模态大模型取得显著进展,但在长而未剪辑的视频中,缺乏用户意图引导的视觉信息易引发冗余计算和噪声干扰。为此,我们提出FocusChat,一种基于文本引导的多模态大语言模型,强调与用户提示相关联的视觉信息。模型首先通过语义提取模块,分别从视觉和文本分支提取图像与文本语义;再通过时空过滤模块(STFM)融合二者,实现显式的空间级信息过滤与隐式的时序特征过滤,确保视觉标记与用户查询高度对齐。该方法大幅减少输入大模型的视觉标记数量。在零样本实验中,FocusChat性能显著超越Video-LLaMA,训练数据量仅为后者的十分之一,仅使用16个视觉标记;在少样本实验中,其表现接近当前最先进水平,仅需0.72M预训练数据。
原文摘要 · Abstract (English)
Recently, multi-modal large language models have made significant progress. However, visual information lacking of guidance from the user's intention may lead to redundant computation and involve unnecessary visual noise, especially in long, untrimmed videos. To address this issue, we propose FocusChat, a text-guided multi-modal large language model (LLM) that emphasizes visual information correlated to the user's prompt. In detail, Our model first undergoes the semantic extraction module, which comprises a visual semantic branch and a text semantic branch to extract image and text semantics, respectively. The two branches are combined using the Spatial-Temporal Filtering Module (STFM). STFM enables explicit spatial-level information filtering and implicit temporal-level feature filtering, ensuring that the visual tokens are closely aligned with the user's query. It lowers the essential number of visual tokens inputted into the LLM. FocusChat significantly outperforms Video-LLaMA in zero-shot experiments, using an order of magnitude less training data with only 16 visual tokens occupied. It achieves results comparable to the state-of-the-art in few-shot experiments, with only 0.72M pre-training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。