arXiv:2412.09907cs.CV2024-12被引 2

提出可自适应问题的视觉压缩器,提升长视频理解效率

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs

  • 基于Transformer的视觉压缩器,按问题需求选择性提取关键帧
  • 在InfiniBench上比现有方法准确率高12.3%,内存消耗降低67%
  • 适合需要高效处理长视频的多模态大模型研究者

随着视频数据复杂度增加,现有长时视频理解方法难以有效捕捉和分析长时间序列。本文提出一种新颖的大规模多模态模型框架,引入名为IQViC(In-context, Question Adaptive Visual Compressor)的视觉压缩器。受人类选择性注意力与上下文记忆机制启发,IQViC采用问题条件化的在上下文压缩策略,而非依赖全视频视觉特征,从而显著减少内存令牌需求。该方法通过Transformer架构实现,仅提取与问题相关的视觉信息。在基于InfiniBench的新数据集及标准基准上的大量实验表明,该框架在视频理解准确率和内存效率方面均优于当前最优方法。

原文摘要 · Abstract (English)

With the increasing complexity of video data and the need for more efficient long-term temporal understanding, existing long-term video understanding methods often fail to accurately capture and analyze extended video sequences. These methods typically struggle to maintain performance over longer durations and to handle the intricate dependencies within the video content. To address these limitations, we propose a simple yet effective large multi-modal model framework for long-term video understanding that incorporates a novel visual compressor, the In-context, Question Adaptive Visual Compressor (IQViC). The key idea, inspired by humans' selective attention and in-context memory mechanisms, is to introduce a novel visual compressor and incorporate efficient memory management techniques to enhance long-term video question answering. Our framework utilizes IQViC, a transformer-based visual compressor, enabling question-conditioned in-context compression, unlike existing methods that rely on full video visual features. This selectively extracts relevant information, significantly reducing memory token requirements. Through extensive experiments on a new dataset based on InfiniBench for long-term video understanding, and standard benchmarks used for existing methods' evaluation, we demonstrate the effectiveness of our proposed IQViC framework and its superiority over state-of-the-art methods in terms of video understanding accuracy and memory efficiency.

视频理解视觉压缩多模态长时记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。