用视频自监督片段微调,提升大模型对视频细节的理解能力。
SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding
- 通过自监督视频片段任务微调,无需人工标注
- 在细粒度视频理解任务上显著提升模型表现
- 适合需要精准视频分析的科研与应用开发者
近年来,基于视频的大语言模型(Video-LLMs)在多模态大模型推动下取得显著进展。尽管这些模型能较好描述视频整体内容,但在视觉动态和视频细节理解方面仍存在不足。本文发现,在自监督片段任务上微调视频大模型,可显著增强其细粒度理解能力。为此提出两项贡献:(1) 自监督片段微调(SF²T),一种无需人工标注、利用视频内在特征进行训练的新方法,有效避免自然语言难以捕捉复杂时空变化的问题;(2) 新建基准数据集 FineVidBench,用于在场景级和片段级双重评估视频大模型性能。实验验证了 SF²T 在多个模型上的有效性,结果表明该方法显著提升了模型对时空细节的捕捉与解析能力。
原文摘要 · Abstract (English)
Video-based Large Language Models (Video-LLMs) have witnessed substantial advancements in recent years, propelled by the advancement in multi-modal LLMs. Although these models have demonstrated proficiency in providing the overall description of videos, they struggle with fine-grained understanding, particularly in aspects such as visual dynamics and video details inquiries. To tackle these shortcomings, we find that fine-tuning Video-LLMs on self-supervised fragment tasks, greatly improve their fine-grained video understanding abilities. Hence we propose two key contributions:(1) Self-Supervised Fragment Fine-Tuning (SF$^2$T), a novel effortless fine-tuning method, employs the rich inherent characteristics of videos for training, while unlocking more fine-grained understanding ability of Video-LLMs. Moreover, it relieves researchers from labor-intensive annotations and smartly circumvents the limitations of natural language, which often fails to capture the complex spatiotemporal variations in videos; (2) A novel benchmark dataset, namely FineVidBench, for rigorously assessing Video-LLMs' performance at both the scene and fragment levels, offering a comprehensive evaluation of their capabilities. We assessed multiple models and validated the effectiveness of SF$^2$T on them. Experimental results reveal that our approach improves their ability to capture and interpret spatiotemporal details.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。