视频微调提升视频理解,但会损害静态图像能力。
Temporal Gains, Spatial Costs: Revisiting Video Fine-Tuning in Multimodal Large Language Models
- 通过自适应帧采样策略,动态分配视频帧数以平衡时空理解。
- 增加采样帧数能提升视频表现,但对图像任务帮助有限甚至有害。
- 适用于需兼顾图像与视频理解的多模态模型优化场景。
多模态大语言模型通常采用多阶段训练,其中基于视频的监督微调(Video-SFT)是提升视觉理解的关键步骤。然而,其对细粒度视觉能力演变的影响,尤其是空间与时间理解之间的平衡,仍不清晰。本文系统研究了Video-SFT如何重塑多模态大模型的视觉能力。在不同架构、参数量和帧采样设置下,我们观察到一致模式:Video-SFT显著提升视频性能,但对静态图像基准测试的改进有限,甚至导致退化。进一步分析表明,该权衡与时间预算密切相关:增加采样帧数可提升视频表现,却无法可靠改善图像性能。为此,我们提出一种指令感知的混合帧策略,能自适应分配帧数,部分缓解图像-视频性能冲突。结果表明,Video-SFT并非免费午餐,保留空间理解仍是图像-视频联合训练的核心挑战。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) are typically trained in multiple stages, with video-based supervised fine-tuning (Video-SFT) serving as a key step for improving visual understanding. Yet its effect on the fine-grained evolution of visual capabilities, particularly the balance between spatial and temporal understanding, remains poorly understood. In this paper, we systematically study how Video-SFT reshapes visual capabilities in MLLMs. Across architectures, parameter scales, and frame sampling settings, we observe a consistent pattern: Video-SFT reliably improves video performance, but often yields limited gains or even degradation on static image benchmarks. We further show that this trade-off is closely tied to temporal budget: increasing the number of sampled frames generally improves video performance, but does not reliably improve static image performance. Motivated by this finding, we study an instruction-aware Hybrid-Frame strategy that adaptively allocates frame counts and partially mitigates the image-video trade-off. Our results indicate that Video-SFT is not a free lunch for MLLMs, and preserving spatial understanding remains a central challenge in joint image-video training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。