arXiv:2504.07745cs.CVcs.AI2025-04CVPR被引 31

用视频自监督片段微调,提升大模型对视频细节的理解能力。

SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding

  • 通过自监督视频片段任务微调,无需人工标注
  • 在细粒度视频理解任务上显著提升模型表现
  • 适合需要精准视频分析的科研与应用开发者

近年来,基于视频的大语言模型(Video-LLMs)在多模态大模型推动下取得显著进展。尽管这些模型能较好描述视频整体内容,但在视觉动态和视频细节理解方面仍存在不足。本文发现,在自监督片段任务上微调视频大模型,可显著增强其细粒度理解能力。为此提出两项贡献:(1) 自监督片段微调(SF²T),一种无需人工标注、利用视频内在特征进行训练的新方法,有效避免自然语言难以捕捉复杂时空变化的问题;(2) 新建基准数据集 FineVidBench,用于在场景级和片段级双重评估视频大模型性能。实验验证了 SF²T 在多个模型上的有效性,结果表明该方法显著提升了模型对时空细节的捕捉与解析能力。

原文摘要 · Abstract (English)

Video-based Large Language Models (Video-LLMs) have witnessed substantial advancements in recent years, propelled by the advancement in multi-modal LLMs. Although these models have demonstrated proficiency in providing the overall description of videos, they struggle with fine-grained understanding, particularly in aspects such as visual dynamics and video details inquiries. To tackle these shortcomings, we find that fine-tuning Video-LLMs on self-supervised fragment tasks, greatly improve their fine-grained video understanding abilities. Hence we propose two key contributions:(1) Self-Supervised Fragment Fine-Tuning (SF$^2$T), a novel effortless fine-tuning method, employs the rich inherent characteristics of videos for training, while unlocking more fine-grained understanding ability of Video-LLMs. Moreover, it relieves researchers from labor-intensive annotations and smartly circumvents the limitations of natural language, which often fails to capture the complex spatiotemporal variations in videos; (2) A novel benchmark dataset, namely FineVidBench, for rigorously assessing Video-LLMs' performance at both the scene and fragment levels, offering a comprehensive evaluation of their capabilities. We assessed multiple models and validated the effectiveness of SF$^2$T on them. Experimental results reveal that our approach improves their ability to capture and interpret spatiotemporal details.

视频理解自监督学习大模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。