arXiv:2409.18938cs.CVcs.AI2024-09被引 32

系统梳理多模态大模型从图像到长视频理解的演进与挑战

From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

  • 对比图像、短视频与长视频任务差异,剖析长视频理解难点
  • 总结针对长视频设计的模型架构与训练方法进展
  • 适合关注视频理解前沿的研究者与开发者参考

将大语言模型(LLMs)与视觉编码器结合,在视觉理解任务中展现出良好性能,能够利用其生成和理解人类语言的能力进行视觉推理。由于视觉数据多样性,多模态大语言模型(MM-LLMs)在图像、短视频和长视频理解任务中存在不同的模型设计与训练方式。本文聚焦于长视频理解相较于静态图像和短视频理解所面临的显著差异与独特挑战:静态图像为单帧空间信息,短视频包含帧间时序与事件内时序,而长视频则包含多事件间的跨事件时序与长期依赖关系。本综述旨在追踪并总结从图像理解到长视频理解的MM-LLM发展脉络,分析各类视觉理解任务之间的区别,强调长视频理解中的细粒度时空细节、动态事件及长期依赖等挑战,并详细梳理了针对长视频理解的模型设计与训练方法进展。最后,我们在不同长度视频理解基准上对比现有MM-LLMs的表现,并讨论其未来发展方向。

原文摘要 · Abstract (English)

The integration of Large Language Models (LLMs) with visual encoders has recently shown promising performance in visual understanding tasks, leveraging their inherent capability to comprehend and generate human-like text for visual reasoning. Given the diverse nature of visual data, MultiModal Large Language Models (MM-LLMs) exhibit variations in model designing and training for understanding images, short videos, and long videos. Our paper focuses on the substantial differences and unique challenges posed by long video understanding compared to static image and short video understanding. Unlike static images, short videos encompass sequential frames with both spatial and within-event temporal information, while long videos consist of multiple events with between-event and long-term temporal information. In this survey, we aim to trace and summarize the advancements of MM-LLMs from image understanding to long video understanding. We review the differences among various visual understanding tasks and highlight the challenges in long video understanding, including more fine-grained spatiotemporal details, dynamic events, and long-term dependencies. We then provide a detailed summary of the advancements in MM-LLMs in terms of model design and training methodologies for understanding long videos. Finally, we compare the performance of existing MM-LLMs on video understanding benchmarks of various lengths and discuss potential future directions for MM-LLMs in long video understanding.

多模态视频理解大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。