arXiv:2505.07344cs.CVcs.AI2025-05NeurIPS被引 18

将扩散模型与自回归机制结合,实现连续潜空间的高质量长视频生成。

Generative Pre-trained Autoregressive Diffusion Transformer

  • 在连续潜空间中自回归预测未来帧,用扩散损失建模运动与语义一致性。
  • 在视频生成、表征能力与少样本学习上均表现优异,优于现有方法。
  • 轻量级因果注意力与无参数时间条件机制,提升训练推理效率。

本文提出GPDiT,一种生成式预训练自回归扩散Transformer,统一扩散模型与自回归建模优势,用于长视频合成,基于连续潜空间。不同于传统离散标记预测,GPDiT自回归地利用扩散损失预测未来潜变量帧,自然建模运动动态与跨帧语义一致性。该连续自回归框架不仅提升生成质量,还赋予模型表征能力。我们引入轻量级因果注意力变体与无参数旋转式时间条件机制,显著提升训练与推理效率。大量实验表明,GPDiT在视频生成质量、视频表征能力及少样本学习任务中表现卓越,展现了其在连续空间中进行视频建模的强大潜力。

原文摘要 · Abstract (English)

In this work, we present GPDiT, a Generative Pre-trained Autoregressive Diffusion Transformer that unifies the strengths of diffusion and autoregressive modeling for long-range video synthesis, within a continuous latent space. Instead of predicting discrete tokens, GPDiT autoregressively predicts future latent frames using a diffusion loss, enabling natural modeling of motion dynamics and semantic consistency across frames. This continuous autoregressive framework not only enhances generation quality but also endows the model with representation capabilities. Additionally, we introduce a lightweight causal attention variant and a parameter-free rotation-based time-conditioning mechanism, improving both the training and inference efficiency. Extensive experiments demonstrate that GPDiT achieves strong performance in video generation quality, video representation ability, and few-shot learning tasks, highlighting its potential as an effective framework for video modeling in continuous space.

视频生成扩散模型自回归连续潜空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。