arXiv:2512.21004cs.CV2025-12被引 1

用预测下一帧训练视频模型,提升视觉表征能力

Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations

  • 通过掩码下一帧预测构建自回归视频预训练框架
  • 在多个下游分类任务中优于现有生成式预训练方法
  • 适合需要高质量视觉表征的视频理解研究者

近期预训练通用基础模型的进展显著提升了多样下游任务的表现。尽管自回归生成模型如GPT已革新自然语言处理,多数视觉生成预训练方法仍依赖BERT风格的掩码建模,常忽略视频分析中至关重要的时序信息。现有少数自回归视觉预训练方法存在语义定位不准、生成质量差等问题。本文提出NExT-Vid,一种新型自回归视觉生成预训练框架,利用掩码下一帧预测联合建模图像与视频。NExT-Vid引入上下文隔离的自回归预测器,解耦语义表示与目标解码;采用条件流匹配解码器,提升生成质量与多样性。通过上下文隔离的流匹配预训练,该方法获得强视觉表征。大规模预训练模型的大量实验表明,本方法在下游分类任务中通过注意力探测一致优于以往生成式预训练方法。

原文摘要 · Abstract (English)

Recent advances in pretraining general foundation models have significantly improved performance across diverse downstream tasks. While autoregressive (AR) generative models like GPT have revolutionized NLP, most visual generative pretraining methods still rely on BERT-style masked modeling, which often disregards the temporal information essential for video analysis. The few existing autoregressive visual pretraining methods suffer from issues such as inaccurate semantic localization and poor generation quality, leading to poor semantics. In this work, we propose NExT-Vid, a novel autoregressive visual generative pretraining framework that utilizes masked next-frame prediction to jointly model images and videos. NExT-Vid introduces a context-isolated autoregressive predictor to decouple semantic representation from target decoding, and a conditioned flow-matching decoder to enhance generation quality and diversity. Through context-isolated flow-matching pretraining, our approach achieves strong representations. Extensive experiments on large-scale pretrained models demonstrate that our proposed method consistently outperforms previous generative pretraining methods for visual representation learning via attentive probing in downstream classification.

视频生成自回归模型表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。