arXiv:2412.10925cs.CVcs.AI2024-12被引 17

用正则化防崩溃,让视频表征学得更抽象有用

Video Representation Learning with Joint-Embedding Predictive Architectures

  • 设计联合嵌入预测架构,加方差协方差正则避免表征坍缩
  • 在需理解物体运动规律的任务上,优于生成基线模型
  • 引入潜在变量捕捉不确定性,适合非确定性视频建模

视频表征学习是机器学习中的重要方向。我们提出带方差-协方差正则化的视频联合嵌入预测架构(VJ-VCR),一种自监督视频表征学习方法,通过方差和协方差正则化防止表征坍缩。结果表明,该模型的隐藏表征包含输入数据的抽象高层信息,在需要理解视频中运动物体内在动态的下游任务中,表现优于生成基线模型。此外,我们探索了在VJ-VCR框架中融入潜在变量的不同方式,以在非确定性场景下捕捉未来状态的不确定性。

原文摘要 · Abstract (English)

Video representation learning is an increasingly important topic in machine learning research. We present Video JEPA with Variance-Covariance Regularization (VJ-VCR): a joint-embedding predictive architecture for self-supervised video representation learning that employs variance and covariance regularization to avoid representation collapse. We show that hidden representations from our VJ-VCR contain abstract, high-level information about the input data. Specifically, they outperform representations obtained from a generative baseline on downstream tasks that require understanding of the underlying dynamics of moving objects in the videos. Additionally, we explore different ways to incorporate latent variables into the VJ-VCR framework that capture information about uncertainty in the future in non-deterministic settings.

视频表征自监督学习联合嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。