提出视频理解新框架Apollo,让大模型高效处理长视频并显著提升性能。
Apollo: An Exploration of Video Understanding in Large Multimodal Models
- 发现规模一致性规律,小模型设计可迁移至大模型,降低研发成本。
- 实验证明采样帧率优于均匀采样,优选视觉编码器提升视频表征能力。
- 适用于需要长视频理解的AI研发与产品团队,尤其关注多模态模型优化者。
尽管大型多模态模型(LMMs)已快速集成视频感知能力,但其视频理解机制仍不清晰,导致诸多设计决策缺乏依据。训练与评估高成本及研究资源有限,阻碍了视频-LMM的发展。为此,本文开展全面研究,揭示驱动视频理解的关键因素。我们发现‘规模一致性’现象:在一定临界规模前,小模型的设计与训练策略可有效迁移至大模型。基于此,我们系统探索了视频采样、架构、数据构成、训练调度等关键问题。例如,证明训练时按帧率采样优于均匀采样,且特定视觉编码器对视频表示更优。据此,我们推出Apollo系列模型,在不同规模下均达顶尖水平。其中,Apollo-3B在LongVideoBench上取得55.1分,超越多数7B模型;Apollo-7B在MLVU上达70.9分,视频-MME上达63.3分,为7B级模型当前最优表现。
原文摘要 · Abstract (English)
Despite the rapid integration of video perception capabilities into Large Multimodal Models (LMMs), the underlying mechanisms driving their video understanding remain poorly understood. Consequently, many design decisions in this domain are made without proper justification or analysis. The high computational cost of training and evaluating such models, coupled with limited open research, hinders the development of video-LMMs. To address this, we present a comprehensive study that helps uncover what effectively drives video understanding in LMMs. We begin by critically examining the primary contributors to the high computational requirements associated with video-LMM research and discover Scaling Consistency, wherein design and training decisions made on smaller models and datasets (up to a critical size) effectively transfer to larger models. Leveraging these insights, we explored many video-specific aspects of video-LMMs, including video sampling, architectures, data composition, training schedules, and more. For example, we demonstrated that fps sampling during training is vastly preferable to uniform frame sampling and which vision encoders are the best for video representation. Guided by these findings, we introduce Apollo, a state-of-the-art family of LMMs that achieve superior performance across different model sizes. Our models can perceive hour-long videos efficiently, with Apollo-3B outperforming most existing $7$B models with an impressive 55.1 on LongVideoBench. Apollo-7B is state-of-the-art compared to 7B LMMs with a 70.9 on MLVU, and 63.3 on Video-MME.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。