arXiv:2606.25473cs.CVcs.LG2026-06被引 7

提出统一教学与自教学的因果扩散蒸馏方法,实现高效视频生成。

Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models

论文配图:Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models
图 1 · 摘自论文原文
  • 融合教师强制与自教学机制,构建因果扩散蒸馏新框架
  • 2步采样下达到84.63的VBench-T2V分数,速度比离散时间快10倍
  • 适用于流式视频生成与交互式世界模型,仅用合成数据训练

自回归视频扩散模型结合因果扩散变换器,已成为实时流式视频生成与动作条件交互世界模型的主要范式。本文将先进的扩散蒸馏框架rCM扩展至自回归视频扩散领域。rCM的核心思想在于正向与反向发散的互补性:一致性模型(CMs)代表正向发散,分布匹配蒸馏(DMD)代表反向发散。该思想自然延伸至自回归设置中,教师强制(TF)提供离线正向发散训练,自教学(SF)则对应在线策略反向发散优化。贡献包括:(1) 实验表明教师强制CM是自教学DMD的最佳初始化策略;(2) 首次实现基于教师强制的连续时间CM(如sCM/MeanFlow)在自回归视频扩散中的应用,得益于自研掩码FlashAttention-2 JVP内核,收敛速度比离散时间CM快10倍;(3) 提出Causal-rCM,一种领先、统一且可扩展的扩散蒸馏与因果训练开源方案;(4) 仅使用合成数据,在帧级与块级设置下均取得当前最优流式视频生成性能。特别地,所蒸馏的2步因果Wan2.1-1.3B模型在仅1或2次采样下即可达84.63的VBench-T2V得分。进一步将Causal-rCM应用于Cosmos 3,一个具备动作条件生成能力的多模态世界基础模型,实现了交互式世界模型。

原文摘要 · Abstract (English)

Autoregressive video diffusion with causal diffusion transformers has emerged as a major paradigm for real-time streaming video generation and action-conditioned interactive world models. In this work, we extend rCM, an advanced diffusion distillation framework, to autoregressive video diffusion. The core philosophy of rCM lies in the complementarity between forward and reverse divergences, represented by consistency models (CMs) and distribution matching distillation (DMD), respectively, in diffusion distillation. This philosophy naturally carries over to the autoregressive setting, where teacher-forcing (TF) provides an offline, forward-divergence causal training paradigm, while self-forcing (SF) corresponds to an on-policy, reverse-divergence refinement. Our contributions are: (1) through extensive experiments, we show that teacher-forcing CM is currently the best complement to self-forcing DMD as an initialization strategy (2) we present the first implementation of teacher-forcing-based continuous-time CMs (e.g., sCM/MeanFlow) for autoregressive video diffusion, enabled by our custom-mask FlashAttention-2 JVP kernel, achieving 10$\times$ faster convergence compared to discrete-time CMs (dCMs) (3) we introduce Causal-rCM, a leading, unified, and scalable algorithm-infrastructure open recipe for diffusion distillation and causal training (4) we achieve state-of-the-art streaming video generation performance in both frame-wise and chunk-wise settings, using only synthetic data for training. Notably, our distilled 2-step causal Wan2.1-1.3B model achieves a VBench-T2V score of 84.63 with only 1 or 2 sampling steps. We further apply Causal-rCM to Cosmos 3, an advanced omnimodal world foundation model for physical AI with action-conditioned generation capability, enabling an interactive world model.

视频生成扩散模型因果推理自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。