arXiv:2412.15212cs.CVcs.AI2024-12被引 32

大模型视频自监督学习可有效提升4D视觉任务性能。

Scaling 4D Representations

  • 用大规模视频数据训练掩码自编码器,构建4D空间时间表示。
  • 模型参数从20M增至220亿,各项4D任务性能持续提升。
  • 适合关注视频理解、3D重建与运动估计的研究者。

自监督视频学习的可扩展性尚未得到充分验证。以往研究多聚焦于语义任务(如动作分类、ImageNet分类),本文转向非语义的时空任务:相机位姿估计、点/物体跟踪和深度估计。实验表明,通过在超大规模视频数据上训练,基于Transformer的掩码自编码器(MAE)模型随规模增长,性能在这些4D任务上持续提升,最大模型达220亿参数,为迄今最大的自监督视频模型。与多种近期图像及视频模型进行严格对比,验证了4D表示的可扩展优势。预训练模型已开源于https://github.com/google-deepmind/representations4d。

原文摘要 · Abstract (English)

Scaling has not yet been convincingly demonstrated for pure self-supervised learning from video. However, prior work has focused evaluations on semantic-related tasks $\unicode{x2013}$ action classification, ImageNet classification, etc. In this paper we focus on evaluating self-supervised learning on non-semantic vision tasks that are more spatial (3D) and temporal (+1D = 4D), such as camera pose estimation, point and object tracking, and depth estimation. We show that by learning from very large video datasets, masked auto-encoding (MAE) with transformer video models actually scales, consistently improving performance on these 4D tasks, as model size increases from 20M all the way to the largest by far reported self-supervised video model $\unicode{x2013}$ 22B parameters. Rigorous apples-to-apples comparison with many recent image and video models demonstrates the benefits of scaling 4D representations. Pretrained models are available at https://github.com/google-deepmind/representations4d .

视频理解自监督4D表示大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。