提出可插拔的显式时间建模模块,提升视频理解模型性能。
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
- 设计可堆叠时间编码器,灵活控制时间感受野与帧压缩率。
- 实验证明显式时间建模显著优于隐式建模,尤其在时序理解任务中。
- 模块通用性强,适用于视频和图像任务,适合改进现有多模态大模型。
将多模态大语言模型(MLLMs)应用于视频理解面临建模跨帧时序关系的重大挑战。现有方法或采用仅依赖大语言模型解码器的隐式时间建模,或使用辅助时间编码器的显式时间建模。为探究两种范式的优劣,我们提出可堆叠时间编码器(STE),支持灵活的显式时间建模,具备可调的时间感受野和标记压缩比。基于STE,我们在整体性能、标记压缩效率及特定时序理解能力等方面系统比较了隐式与显式建模。同时研究了STE的设计考量及其作为通用插件模块在视频与图像模态中的影响。结果强调了显式时间建模的关键作用,为推进视频MLLMs提供了切实可行的见解。
原文摘要 · Abstract (English)
Applying Multimodal Large Language Models (MLLMs) to video understanding presents significant challenges due to the need to model temporal relations across frames. Existing approaches adopt either implicit temporal modeling, relying solely on the LLM decoder, or explicit temporal modeling, employing auxiliary temporal encoders. To investigate this debate between the two paradigms, we propose the Stackable Temporal Encoder (STE). STE enables flexible explicit temporal modeling with adjustable temporal receptive fields and token compression ratios. Using STE, we systematically compare implicit and explicit temporal modeling across dimensions such as overall performance, token compression effectiveness, and temporal-specific understanding. We also explore STE's design considerations and broader impacts as a plug-in module and in image modalities. Our findings emphasize the critical role of explicit temporal modeling, providing actionable insights to advance video MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。