arXiv:2410.11417cs.CVcs.MM2024-10被引 6

通过记忆增强压缩,提升大模型对视频时序关系的理解能力

VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models

  • 设计双压缩机制,融合短时与长时视频依赖关系
  • 在多个视频问答数据集上显著超越现有模型性能
  • 适合需要精细时序理解的视频分析任务

基于视频的多模态大语言模型(Video-LLMs)在视频理解任务中具有巨大潜力。然而,多数现有模型将视频视为独立帧的序列,导致时空交互不足,难以实现细粒度理解,且受限于视觉标记容量,难以处理长视频。为此,我们提出VidCompress,一种具备记忆增强时序压缩能力的新视频大模型。该模型采用双压缩架构:记忆增强压缩器利用多尺度变换器与记忆缓存机制,捕捉视频中的短期与长期时序关系,并压缩视觉标记;文本感知压缩器则通过Q-Former结合时序上下文到查询嵌入中,生成紧凑的视觉标记。在多个视频问答数据集和综合基准测试中,VidCompress均展现出高效建模复杂时空关系的能力,并显著优于现有视频大模型。

原文摘要 · Abstract (English)

Video-based multimodal large language models (Video-LLMs) possess significant potential for video understanding tasks. However, most Video-LLMs treat videos as a sequential set of individual frames, which results in insufficient temporal-spatial interaction that hinders fine-grained comprehension and difficulty in processing longer videos due to limited visual token capacity. To address these challenges, we propose VidCompress, a novel Video-LLM featuring memory-enhanced temporal compression. VidCompress employs a dual-compressor approach: a memory-enhanced compressor captures both short-term and long-term temporal relationships in videos and compresses the visual tokens using a multiscale transformer with a memory-cache mechanism, while a text-perceived compressor generates condensed visual tokens by utilizing Q-Former and integrating temporal contexts into query embeddings with cross attention. Experiments on several VideoQA datasets and comprehensive benchmarks demonstrate that VidCompress efficiently models complex temporal-spatial relations and significantly outperforms existing Video-LLMs.

视频理解时序压缩大模型记忆机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。