无需训练即可快速生成4K高清图像,速度比现有方法快10至35倍。
PixelRush: Ultra-Fast, Training-Free High-Resolution Image Generation via One-step Diffusion
- 采用单步扩散与分块推理,跳过多次迭代重建过程。
- 20秒生成4K图像,较顶尖方法提速10至35倍且画质优异。
- 适合需要高速高分辨率图像生成的实时应用开发者。
预训练扩散模型虽能生成高质量图像,但受限于原始训练分辨率。近期无训练方法尝试通过去噪过程中干预突破此限制,但计算开销大,生成一张4K图像常需超过五分钟。本文提出PixelRush,首个无需调参的实用高分辨率文本到图像生成框架。基于分块推理范式,消除多轮反演与重生成需求,实现低步数下的高效分块去噪。为缓解少步数生成中的拼接伪影,提出无缝融合策略;并通过噪声注入机制减轻过度平滑问题。实验表明,PixelRush在约20秒内生成4K图像,相较最先进方法提速10×至35×,同时保持卓越视觉保真度。
原文摘要 · Abstract (English)
Pre-trained diffusion models excel at generating high-quality images but remain inherently limited by their native training resolution. Recent training-free approaches have attempted to overcome this constraint by introducing interventions during the denoising process; however, these methods incur substantial computational overhead, often requiring more than five minutes to produce a single 4K image. In this paper, we present PixelRush, the first tuning-free framework for practical high-resolution text-to-image generation. Our method builds upon the established patch-based inference paradigm but eliminates the need for multiple inversion and regeneration cycles. Instead, PixelRush enables efficient patch-based denoising within a low-step regime. To address artifacts introduced by patch blending in few-step generation, we propose a seamless blending strategy. Furthermore, we mitigate over-smoothing effects through a noise injection mechanism. PixelRush delivers exceptional efficiency, generating 4K images in approximately 20 seconds representing a 10$\times$ to 35$\times$ speedup over state-of-the-art methods while maintaining superior visual fidelity. Extensive experiments validate both the performance gains and the quality of outputs achieved by our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。