用压缩视频结构提升运动感知,让视频多模态模型更高效
Efficient Motion-Aware Video MLLM
- 用GOP单元融合空间与运动信息,生成紧凑视觉特征
- 在MotionBench上超越现有模型,推理成本降低30%以上
- 适合需要高效处理长视频和运动理解的场景
当前多数视频多模态大模型依赖均匀采样帧和图像级编码器,导致数据处理效率低且运动感知有限。为此,我们提出EMA——一种基于压缩视频结构输入的高效运动感知视频多模态大模型。设计了运动感知的GOP(图像组)编码器,在压缩视频流中融合空间与运动信息,生成紧凑且富含信息的视觉标记。通过在原生慢-快输入架构中整合较少但更密集的RGB帧与较多但更稀疏的运动向量,有效减少冗余并增强运动表征能力。此外,我们构建了MotionBench基准,用于评估四类运动理解:线性、曲线、旋转和接触式运动。实验表明,EMA在MotionBench及主流视频问答基准上均达顶尖性能,同时显著降低推理开销。更重要的是,其在长视频理解任务中表现出优异可扩展性。
原文摘要 · Abstract (English)
Most current video MLLMs rely on uniform frame sampling and image-level encoders, resulting in inefficient data processing and limited motion awareness. To address these challenges, we introduce EMA, an Efficient Motion-Aware video MLLM that utilizes compressed video structures as inputs. We propose a motion-aware GOP (Group of Pictures) encoder that fuses spatial and motion information within a GOP unit in the compressed video stream, generating compact, informative visual tokens. By integrating fewer but denser RGB frames with more but sparser motion vectors in this native slow-fast input architecture, our approach reduces redundancy and enhances motion representation. Additionally, we introduce MotionBench, a benchmark for evaluating motion understanding across four motion types: linear, curved, rotational, and contact-based. Experimental results show that EMA achieves state-of-the-art performance on both MotionBench and popular video question answering benchmarks, while reducing inference costs. Moreover, EMA demonstrates strong scalability, as evidenced by its competitive performance on long video understanding benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。