arXiv:2605.15178cs.CV2026-05被引 28

SANA-WM用2.6亿参数实现分钟级高清视频生成,效率远超同类模型。

SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer

论文配图:SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer
图 1 · 摘自论文原文
  • 混合线性注意力结合门控差分网络与软最大注意力,高效建模长时序画面。
  • 支持6自由度精准相机控制,生成视频动作跟随准确率更高。
  • 可在单张显卡上生成60秒720p视频,适合部署在消费级硬件。

我们提出 SANA-WM,一个2.6B参数的开源世界模型,原生训练用于生成一分钟视频,可合成高保真、720p分辨率、具备精确相机控制的分钟级视频。其视觉质量媲美大型工业级基线模型如 LingBot-World 与 HY-WorldPlay,同时显著提升效率。四个核心设计驱动架构:(1) 混合线性注意力结合帧级门控差分网络(GDN)与软最大注意力,实现内存高效的长序列建模;(2) 双分支相机控制确保6-DoF轨迹精确跟踪;(3) 两阶段生成流程对第一阶段输出使用长视频精炼器,提升序列一致性和质量;(4) 强健的标注流程从公开视频中提取度量尺度的6-DoF相机位姿,生成高质量时空一致的动作标签。基于这些设计,SANA-WM 在数据、训练算力和推理硬件上均表现出卓越效率:仅需约21.3万段公共视频片段与度量尺度位姿监督,64张H100训练15天即可完成,单卡即可生成每段60秒视频;其蒸馏版本经NVFP4量化后可在单张RTX 5090上34秒内去噪生成一段60秒720p视频。在我们的分钟级世界模型基准测试中,SANA-WM 的动作跟随准确率优于先前开源基线,并在36倍更高的吞吐量下达到相近的视觉质量,适用于可扩展的世界建模。

原文摘要 · Abstract (English)

We introduce SANA-WM, an efficient 2.6B-parameter open-source world model natively trained for one-minute generation, synthesizing high-fidelity, 720p, minute-scale videos with precise camera control. SANA-WM achieves visual quality comparable to large-scale industrial baselines such as LingBot-World and HY-WorldPlay, while significantly improving efficiency. Four core designs drive our architecture: (1) Hybrid Linear Attention combines frame-wise Gated DeltaNet (GDN) with softmax attention for memory-efficient long-context modeling. (2) Dual-Branch Camera Control ensures precise 6-DoF trajectory adherence. (3) Two-Stage Generation Pipeline applies a long-video refiner to stage-1 outputs, improving quality and consistency across sequences. (4) Robust Annotation Pipeline extracts accurate metric-scale 6-DoF camera poses from public videos to yield high-quality, spatiotemporally consistent action labels. Driven by these designs, SANA-WMdemonstrates remarkable efficiency across data, training compute, and inference hardware: it uses only $\sim$213K public video clips with metric-scale pose supervision, completes training in 15 days on 64 H100s, and generates each 60s clip on a single GPU; its distilled variant can be deployed on a single RTX 5090 with NVFP4 quantization to denoise a 60s 720p clip in 34s. On our one-minute world-model benchmark, SANA-WM demonstrates stronger action-following accuracy than prior open-source baselines and achieves comparable visual quality at $36\times$ higher throughput for scalable world modeling.

世界模型扩散模型视频生成高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。