arXiv:2602.17869cs.CV2026-02被引 4

为长视频理解设计高效压缩与采样方法,提升大模型处理能力

Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models

  • 基于信息密度自适应采样+时空压缩编码,端到端优化长视频输入
  • 在多个基准上实现高压缩率下仍保持优异的视频理解性能
  • 适合需要处理长时间视频的大模型应用,如智能监控、影视分析

随着视频骨干网络的进步和大语言模型(LLMs)的突破,长达数十分钟的长视频分析已成为可能且日益普遍。然而,视频序列固有的冗余性给当前先进模型带来双重挑战:一是受限于内存,在有限条件下纳入更多帧;二是从海量输入中提取有判别性的信息。本文提出一种全新的长视频理解端到端框架,包含基于信息密度的自适应视频采样器(AVS)和基于自动编码器的时空视频压缩器(SVC),并与多模态大语言模型(MLLM)集成。该系统可自适应地从不同长度的视频序列中高效捕获关键信息,并实现高压缩率的同时保留重要判别特征。在多个基准测试中表现优异,不仅在长视频理解任务上领先,也在标准视频理解任务中展现出强大泛化能力,充分验证了其在处理长时视频复杂性方面的有效性与通用性。

原文摘要 · Abstract (English)

With recent advancements in video backbone architectures, combined with the remarkable achievements of large language models (LLMs), the analysis of long-form videos spanning tens of minutes has become both feasible and increasingly prevalent. However, the inherently redundant nature of video sequences poses significant challenges for contemporary state-of-the-art models. These challenges stem from two primary aspects: 1) efficiently incorporating a larger number of frames within memory constraints, and 2) extracting discriminative information from the vast volume of input data. In this paper, we introduce a novel end-to-end schema for long-form video understanding, which includes an information-density-based adaptive video sampler (AVS) and an autoencoder-based spatiotemporal video compressor (SVC) integrated with a multimodal large language model (MLLM). Our proposed system offers two major advantages: it adaptively and effectively captures essential information from video sequences of varying durations, and it achieves high compression rates while preserving crucial discriminative information. The proposed framework demonstrates promising performance across various benchmarks, excelling in both long-form video understanding tasks and standard video understanding benchmarks. These results underscore the versatility and efficacy of our approach, particularly in managing the complexities of prolonged video sequences.

长视频理解视频压缩多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。