arXiv:2606.14765cs.CVcs.AI2026-06

通过预测未来隐变量,实现无需标签的视频表征学习。

Momentum-Guided Semantic Forecasting (MoFore) for Self-Supervised Video Representation Learning

论文配图:Momentum-Guided Semantic Forecasting (MoFore) for Self-Supervised Video Representation Learning
图 1 · 摘自论文原文
  • 用远处上下文预测未来隐向量,学习时序预测表征
  • 在UCF101上实现强时序稳定性与类别级结构
  • 适合关注视频时序建模与自监督学习的研究者

自监督视频表征学习近年来通过对比学习、掩码重建和预测表征学习取得进展。基于重建的方法如MAE和VideoMAE通过恢复被遮蔽的视觉内容学习表征,而对比方法如CLIP则通过表征对齐学习语义有意义的嵌入空间。本文提出一种动量引导的语义预测框架(MoFore),不依赖像素级重建或特定任务的语义对齐,而是通过从远距离上下文片段预测未来隐向量来学习时序可预测的视频表征。为提升跨时序尺度的鲁棒性,训练中引入随机时序间隔预测。该框架结合预测隐向量与对比正则化,以增强时序一致性并防止表示崩溃。在UCF101数据集上的实验表明,所提框架在无动作标签情况下学习到时序一致且语义有意义的视频表征。定量分析显示嵌入空间具有强时序稳定性与涌现的类别级结构,定性检索实验揭示相关活动间具有运动感知的组织特性。结果表明,长程隐向量预测是一种有效且计算高效的自监督视频表征学习方法,无需依赖重建目标。

原文摘要 · Abstract (English)

Self-supervised video representation learning has recently advanced through contrastive learning, masked reconstruction, and predictive representation learning. Reconstruction-based approaches such as MAE and VideoMAE learn representations by recovering masked visual content \cite{he2022mae,tong2022videomae}, while contrastive methods such as CLIP learn semantically meaningful embedding spaces through representation alignment \cite{radford2021clip}. In this work, we introduce a Momentum-Guided Semantic Forecasting framework (MoFore) for self-supervised video representation learning. Instead of optimizing for pixel-level reconstruction or task-specific semantic alignment, the proposed method learns temporally predictive video representations by forecasting future latent embeddings from temporally distant context clips. To improve robustness across temporal scales, we further introduce randomized temporal-gap forecasting during training. The framework combines predictive latent forecasting with contrastive regularization to encourage temporal consistency while preventing representation collapse. Experiments on the UCF101 dataset demonstrate that the proposed framework learns temporally consistent and semantically meaningful video representations without using action labels during training. Quantitative analysis shows strong temporal stability and emergent category-level structure in the learned embedding space, while qualitative retrieval experiments reveal motion-aware organization across related activities. Overall, the results suggest that long-range latent forecasting provides an effective and computationally efficient approach for self-supervised video representation learning without relying on reconstruction-based objectives.

自监督学习视频表征时序预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。