arXiv:2503.18943cs.CV2025-03被引 36

小模型也能懂长视频,高效理解更省算力。

SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

  • 用双流结构+轻量训练,实现高效视频理解
  • 1B~7B小模型在长视频任务上达顶尖水平
  • 适合移动端部署,兼顾性能与效率

我们提出SlowFast-LLaVA-1.5(SF-LLaVA-1.5),一种面向长视频理解的高效视频大模型家族。通过将双流SlowFast机制融入精简训练流程,并在仅使用公开数据集的混合数据上进行联合视频图像训练,重点优化1B和3B规模的小模型。实验表明,即使规模较小的视频大模型也能在多项视频理解任务中达到领先性能,满足移动设备友好型模型需求。SF-LLaVA-1.5在多种视频与图像任务中表现优异,所有规模(1B至7B)均具备稳健表现。尤其在长视频理解任务(如LongVideoBench和MLVU)中达到当前最优结果,且在小规模下依然保持卓越性能。

原文摘要 · Abstract (English)

We introduce SlowFast-LLaVA-1.5 (abbreviated as SF-LLaVA-1.5), a family of video large language models (LLMs) offering a token-efficient solution for long-form video understanding. We incorporate the two-stream SlowFast mechanism into a streamlined training pipeline, and perform joint video-image training on a carefully curated data mixture of only publicly available datasets. Our primary focus is on highly efficient model scales (1B and 3B), demonstrating that even relatively small Video LLMs can achieve state-of-the-art performance on video understanding, meeting the demand for mobile-friendly models. Experimental results demonstrate that SF-LLaVA-1.5 achieves superior performance on a wide range of video and image tasks, with robust results at all model sizes (ranging from 1B to 7B). Notably, SF-LLaVA-1.5 achieves state-of-the-art results in long-form video understanding (e.g., LongVideoBench and MLVU) and excels at small scales across various video benchmarks.

视频理解小模型高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。