提出PixelWizard框架,实现2K/4K视频高效高保真生成。
PixelWizard: Towards Efficient High-Fidelity Video Generation at Ultra-Large Spatial Resolution

- 分层解耦结构建模与细节生成,用锚点引导高分辨率合成。
- 2K/4K视频生成速度提升10倍以上,保持高质量细节。
- 适合需要超高清视频生成的工业级应用或研究者。
高分辨率视频生成面临优化不稳与计算成本过高的双重瓶颈。密集的标记序列不仅使优化偏向局部纹理而牺牲全局一致性,导致结构坍塌,还带来高昂的训练成本和严重的推理延迟。为此,我们提出PixelWizard框架,通过层级解耦全局结构建模与细粒度细节合成:先建立紧凑的时空锚点以集中密集的结构先验,再以此引导高分辨率下的细节生成。该设计缓解了局部优化偏差,保障结构稳定性而不损失高频细节。在此基础上,引入噪声跨度对齐捷径训练,显式建模步长,使模型能以大步长穿越生成轨迹,突破推理瓶颈。关键的是,结合指数索引偏置采样与自适应噪声跨度校准,使优化与高分辨率网格的偏移噪声调度对齐,实现鲁棒的少步推理,且无需蒸馏带来的沉重开销。大量实验表明,PixelWizard在保持卓越视觉质量的同时,使原生2K/4K视频生成的采样速度提升超过10倍。
原文摘要 · Abstract (English)
High-resolution video generation faces a coupled bottleneck of optimization instability and prohibitive computational costs. The massive expansion of the token sequence not only biases optimization toward local textures at the expense of global coherence, leading to structural collapse, but also imposes prohibitive training costs and severe inference latency. To address this, we propose PixelWizard, a framework that hierarchically decouples global structure modeling from fine-grained detail synthesis. PixelWizard first establishes a compact spatiotemporal anchor to concentrate dense structural priors, which then guides fine-grained generation at high resolution. This mitigates the local optimization bias to ensure structural stability without compromising high-frequency details. Leveraging this structural stability, we introduce Noise-Span Aligned Shortcut Training to break the inference bottleneck. By explicitly modeling the step size, this mechanism allows the model to traverse the generation trajectory with large steps. Crucially, we incorporate Exponential Index-Biased Sampling and Adaptive Noise-Span Calibration to align optimization with the shifted noise schedules of high-resolution grids, ensuring robust few-step inference without incurring the heavy overhead of distillation. Extensive experiments demonstrate that PixelWizard achieves superior visual quality while accelerating the generative sampling of native 2K/4K videos by over 10x.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。