用状态空间模型压缩长视频特征,降低大模型计算开销
State-Space Hierarchical Compression with Gated Attention and Learnable Sampling for Hour-Long Video Understanding in Large Multimodal Models
- 用门控跳跃连接和可学习加权池化实现时空分层压缩
- 在长视频任务中性能接近顶尖模型,令牌数大幅减少
- 适合资源受限场景下大规模视频理解应用
我们提出一种高效框架,在将小时级视频输入大型多模态模型前压缩海量帧特征,缓解因长视频导致的严重令牌爆炸问题。设计采用双向状态空间模型,结合门控跳跃连接与周期性插入可学习查询的可学习加权平均池化机制,实现时空维度的分层下采样,以低成本保持性能。在多个挑战性小时级视频理解任务中,该方法表现媲美当前最优模型,同时显著降低整体令牌预算。值得注意的是,若用传统模块替代状态空间模型,性能出现显著下降,凸显状态空间建模在有效压缩多帧视频信息上的优势。框架强调资源敏感型效率,适用于真实场景部署。我们在多个基准上验证了其可扩展性与通用性,达成高效资源利用与全面视频理解的双重目标。
原文摘要 · Abstract (English)
We propose an efficient framework to compress massive video-frame features before feeding them into large multimodal models, thereby mitigating the severe token explosion arising from hour-long videos. Our design leverages a bidirectional state-space model equipped with a gated skip connection and a learnable weighted-average pooling mechanism applied to periodically inserted learned queries. This structure enables hierarchical downsampling across both spatial and temporal dimensions, preserving performance in a cost-effective manner. Across challenging hour-long video understanding tasks, our approach demonstrates competitive results against state-of-the-art models, while significantly reducing overall token budget. Notably, replacing our state-space model with conventional modules results in substantial performance degradation, highlighting the advantages of the proposed state-space modeling for effectively compressing multi-frame video information. Our framework emphasizes resource-conscious efficiency, making it practical for real-world deployments. We validate its scalability and generality across multiple benchmarks, achieving the dual objectives of efficient resource usage and comprehensive video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。