通过聚类分组保留关键帧,提升视频语言预训练效率与效果
Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
- 按语义聚类分组视觉标记,只保留每组中时间密度最高的标记
- 在高遮蔽率下仍保持完整视频内容,避免时间信息泄露
- 适合追求高效视频多模态模型的研究者和开发者
大规模视频-文本预训练虽具备强大泛化能力,但计算开销巨大。现有掩码视觉建模方法仍存在两大问题:高遮蔽率下视觉信息严重丢失,以及因帧间相关性导致的时间信息泄露。为此,我们提出 ClusterSTM,一种面向高效视频-文本预训练的聚类式时空掩码策略。该方法首先在帧内进行聚类,将视觉标记划分为多个语义独立的簇,再在每个簇中保留时间密度最高的标记。此策略确保保留的标记既能捕捉整体视频内容,又具有强时间相关性。此外,引入视频-文本相关性重建目标,对齐高层多模态语义,超越传统视觉重建。在多个基准上的实验表明,ClusterSTM 在视频-文本检索、视频问答和视频字幕任务上均取得优异表现,成为高效视频-文本模型的新标杆。
原文摘要 · Abstract (English)
Large-scale video-language pretraining enables strong generalization across multimodal tasks but often incurs prohibitive computational costs. Although recent advances in masked visual modeling help mitigate this issue, they still suffer from two fundamental limitations: severe visual information loss under high masking ratios and temporal information leakage caused by inter-frame correlations. To address these challenges, we propose ClusterSTM, a Cluster-Wise Spatio-Temporal Masking strategy for efficient video-language pretraining. ClusterSTM first performs intra-frame clustering to partition visual tokens into multiple semantically independent clusters, then conducts cluster-wise masking by retaining the token with the highest temporal density within each cluster. Our masking strategy ensure that the retained tokens capture holistic video content while exhibit strong temporal correlation. Additionally, we introduce a video-text relevance reconstruction objective that aligns high-level multimodal semantics beyond conventional visual reconstruction. Extensive experiments across multiple benchmarks demonstrate that ClusterSTM achieves superior performance on video-text retrieval, video question answering, and video captioning tasks, establishing a new state-of-the-art among efficient video-language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。