通过解码视觉-语言模型的时间理解机制,提升视频理解性能
Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding
- 发现视觉编码器与语言模型间中间接口是时间理解关键
- 提出面向时间的训练方案与扩展接口,显著提升视频任务表现
- 适合关注视觉-语言模型视频应用的研究者参考
近年来,大型视觉-语言模型(LVLMs)取得显著进展。为应对视频理解任务,多数模型依赖其隐含的时间理解能力,却未揭示支撑该能力的关键组件,可能限制其在视频理解中的潜力。本文开展全面的实证研究,解析影响LVLM时间理解能力的核心因素。结果表明,视觉编码器与大语言模型之间的中间接口起决定性作用。基于此,我们提出一种面向时间的优化方案,包括时间导向的训练策略和扩大的接口结构。使用该方案构建的模型在标准视频理解任务上显著优于现有方法。
原文摘要 · Abstract (English)
Recent years have witnessed outstanding advances of large vision-language models (LVLMs). In order to tackle video understanding, most of them depend upon their implicit temporal understanding capacity. As such, they have not deciphered important components that contribute to temporal understanding ability, which might limit the potential of these LVLMs for video understanding. In this work, we conduct a thorough empirical study to demystify crucial components that influence the temporal understanding of LVLMs. Our empirical study reveals that significant impacts are centered around the intermediate interface between the visual encoder and the large language model. Building on these insights, we propose a temporal-oriented recipe that encompasses temporal-oriented training schemes and an upscaled interface. Our final model developed using our recipe significantly enhances previous LVLMs on standard video understanding tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。