用图文模型迁移能力,低成本实现视频理解,效果接近专用模型。
Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- 将图文基础模型的视觉能力迁移到视频任务,分冻结特征与适配特征两类方法。
- 在时空定位、视频问答等任务上表现媲美甚至超越从零训练的视频模型。
- 适合想快速构建视频理解系统的研究者和开发者,尤其关注高效迁移的场景。
图文基础模型(ILFMs)在视觉-语言理解方面取得显著成功,提供了可迁移的多模态表示,能泛化到多种图像下游任务。视频-文本研究的发展推动了将图像模型扩展至视频领域的兴趣。这种范式称为图像到视频的迁移学习,相比从零训练视频-语言模型,大幅降低了数据与计算需求,同时达到相当或更优的性能。本综述首次全面回顾该新兴领域,首先总结常用ILFMs及其能力。随后根据将图像理解能力迁移至视频任务的范式,系统地将现有技术分为两大类(冻结特征与适配特征),并包含多个细粒度子类别。基于任务特异性,本文详述各类策略,并分析其在从细粒度(如时空视频定位)到粗粒度(如视频问答)视频-文本学习任务中的应用。进一步通过实验分析不同迁移范式的有效性,涵盖多种下游视频理解任务。最后,识别当前挑战并提出未来研究方向。通过提供结构化全景图,本综述旨在为基于现有ILFM推进视频-文本学习建立路线图,激发该快速演进领域的未来研究。
原文摘要 · Abstract (English)
Image-Language Foundation Models (ILFMs) have demonstrated remarkable success in vision-language understanding, providing transferable multimodal representations that generalize across diverse downstream image-based tasks. The advancement of video-text research has spurred growing interest in extending image-based models to the video domain. This paradigm, termed as image-to-video transfer learning, effectively mitigates the substantial data and computational demands compared to training video-language models from scratch while achieves comparable or even stronger model performance. This survey provides the first comprehensive review of this emerging field, which begins by summarizing the widely used ILFMs and their capabilities. We then systematically classify existing image-to-video transfer learning techniques into two broad root categories (frozen features and adapted features), along with numerous fine-grained subcategories, based on the paradigm for transferring image understanding capability to video tasks. Building upon the task-specific nature of image-to-video transfer, this survey methodically elaborates these strategies and details their applications across a spectrum of video-text learning tasks, ranging from fine-grained settings (e.g., spatio-temporal video grounding) to coarse-grained ones (e.g., video question answering). We further present a detailed experimental analysis to investigate the efficacy of different image-to-video transfer learning paradigms on a range of downstream video understanding tasks. Finally, we identify prevailing challenges and highlight promising directions for future research. By offering a comprehensive and structured overview, this survey aims to establish a structured roadmap for advancing video-text learning based on existing ILFM, and to inspire future research directions in this rapidly evolving domain. Github repository is available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。