动态分配计算资源,让视频生成更流畅无闪烁。
Ride the Wave: Precision-Allocated Sparse Attention for Smooth Video Generation

- 根据生成过程的语义变化动态调整计算量,关键帧重点处理。
- 通过分组近似逼近硬件性能,提升计算效率且保持精度。
- 随机选择注意力路径,消除重复抖动,适合高帧率视频生成。
视频扩散变换器已革新高质量视频生成,但自注意力机制带来巨大计算负担。稀疏注意力虽具加速潜力,但现有方法常因固定稀疏模式和确定性块路由引发严重视觉闪烁。为此,我们提出无需训练的精密分配稀疏注意力(PASA)框架,实现高效且时序平滑的视频生成。首先,设计曲率感知的动态预算机制,通过分析生成轨迹在时间步上的加速度,弹性分配精确计算预算,在关键语义转换阶段严格保障高精度处理。其次,以硬件对齐的分组近似替代全局同质估计,有效捕捉细粒度局部变化,同时维持峰值计算吞吐量。最后,将随机选择偏差引入注意力路由机制,采用概率化策略软化刚性选择边界,消除选择振荡,彻底解决导致时间闪烁的局部计算饥饿问题。在主流视频扩散模型上的广泛评估表明,PASA在显著加速推理的同时,持续生成极其流畅、结构稳定的视频序列。
原文摘要 · Abstract (English)
Video Diffusion Transformers have revolutionized high-fidelity video generation but suffer from the massive computational burden of self-attention. While sparse attention provides a promising acceleration solution, existing methods frequently provoke severe visual flickering caused by static sparsity patterns and deterministic block routing. To resolve these limitations, we propose Precision-Allocated Sparse Attention (PASA), a training-free framework designed for highly efficient and temporally smooth video generation. First, we implement a curvature-aware dynamic budgeting mechanism. By profiling the generation trajectory acceleration across timesteps, we elastically allocate the exact-computation budget to secure high-precision processing strictly during critical semantic transitions. Second, we replace global homogenizing estimations with hardware-aligned grouped approximations, successfully capturing fine-grained local variations while maintaining peak compute throughput. Finally, we incorporate a stochastic selection bias into the attention routing mechanism. This probabilistic approach softens rigid selection boundaries and eliminates selection oscillation, effectively eradicating the localized computational starvation that drives temporal flickering. Extensive evaluations on leading video diffusion models demonstrate that PASA achieves substantial inference acceleration while consistently producing remarkably fluid and structurally stable video sequences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。