arXiv:2601.12719cs.CV2026-01被引 5

S2DiT让手机实时生成高质量视频,突破算力瓶颈。

S2DiT: Sandwich Diffusion Transformer for Mobile Streaming Video Generation

  • 采用夹心结构与新型高效注意力机制,提升生成效率。
  • 在iPhone上实现超10帧每秒的流畅视频流,质量媲美服务器模型。
  • 适合移动端视频生成、低延迟应用开发人员使用。

扩散Transformer(DiTs)近期提升了视频生成质量,但其高昂的计算成本使其难以实现实时或设备端生成。本文提出S2DiT,一种面向移动硬件的流式夹心扩散Transformer,可实现高效、高保真、流式视频生成。S2DiT生成更多标记(tokens),但通过创新的高效注意力机制——线性卷积混合注意力(LCHA)与步进自注意力(SSA)保持高效。基于此,我们通过预算感知的动态规划搜索发现最优夹心结构,在质量和效率上均表现优异。此外,提出一种两阶段蒸馏框架,将大型教师模型(如Wan 2.2-14B)的能力迁移到紧凑的少步夹心模型中。最终,S2DiT在保持与最先进服务器视频模型相当质量的同时,可在iPhone上实现超过10 FPS的流式生成。

原文摘要 · Abstract (English)

Diffusion Transformers (DiTs) have recently improved video generation quality. However, their heavy computational cost makes real-time or on-device generation infeasible. In this work, we introduce S2DiT, a Streaming Sandwich Diffusion Transformer designed for efficient, high-fidelity, and streaming video generation on mobile hardware. S2DiT generates more tokens but maintains efficiency with novel efficient attentions: a mixture of LinConv Hybrid Attention (LCHA) and Stride Self-Attention (SSA). Based on this, we uncover the sandwich design via a budget-aware dynamic programming search, achieving superior quality and efficiency. We further propose a 2-in-1 distillation framework that transfers the capacity of large teacher models (e.g., Wan 2.2-14B) to the compact few-step sandwich model. Together, S2DiT achieves quality on par with state-of-the-art server video models, while streaming at over 10 FPS on an iPhone.

视频生成扩散模型移动端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。