综述多模态大模型在视频时序定位中的应用进展
A Survey on Video Temporal Grounding with Multimodal Large Language Model
- 从模型功能、训练范式、特征处理三维度梳理技术体系
- 揭示多模态大模型在零样本、跨任务场景下的强泛化能力
- 适合关注视频理解与大模型融合的科研人员参考
近年来,视频时序定位(VTG)的进展显著提升了细粒度视频理解能力,主要得益于多模态大语言模型(MLLMs)的发展。基于MLLM的VTG方法(VTG-MLLMs)凭借卓越的多模态理解和推理能力,正逐步超越传统微调方法,在零样本、多任务和跨领域设置中表现优异。尽管已有大量关于通用视频-语言理解的综述,但针对VTG-MLLMs的系统性回顾仍较缺乏。为此,本综述通过三维分类体系系统分析当前研究:1)MLLM的功能角色,突出其架构重要性;2)训练范式,分析时间推理与任务适配策略;3)视频特征处理技术,决定时空表征效果。进一步讨论了基准数据集、评估协议,并总结实证发现。最后,指出现存局限并提出未来研究方向。更多资源详见:https://github.com/ki-lw/Awesome-MLLMs-for-Video-Temporal-Grounding。
原文摘要 · Abstract (English)
The recent advancement in video temporal grounding (VTG) has significantly enhanced fine-grained video understanding, primarily driven by multimodal large language models (MLLMs). With superior multimodal comprehension and reasoning abilities, VTG approaches based on MLLMs (VTG-MLLMs) are gradually surpassing traditional fine-tuned methods. They not only achieve competitive performance but also excel in generalization across zero-shot, multi-task, and multi-domain settings. Despite extensive surveys on general video-language understanding, comprehensive reviews specifically addressing VTG-MLLMs remain scarce. To fill this gap, this survey systematically examines current research on VTG-MLLMs through a three-dimensional taxonomy: 1) the functional roles of MLLMs, highlighting their architectural significance; 2) training paradigms, analyzing strategies for temporal reasoning and task adaptation; and 3) video feature processing techniques, which determine spatiotemporal representation effectiveness. We further discuss benchmark datasets, evaluation protocols, and summarize empirical findings. Finally, we identify existing limitations and propose promising research directions. For additional resources and details, readers are encouraged to visit our repository at https://github.com/ki-lw/Awesome-MLLMs-for-Video-Temporal-Grounding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。