无需训练,生成与高分辨率图像感知一致的低分辨率预览图,提升扩散模型效率。
Training-free, Perceptually Consistent Low-Resolution Previews with High-Resolution Image for Efficient Workflows of Diffusion Models
- 基于对易零条件设计无训练方案,确保低分辨与高分辨图像感知一致。
- 计算量减少33%,结合加速技术最高可提速3倍。
- 适用于图像变形、平移等操作,通用性强,适合快速筛选高质量图像。
图像生成模型已成为大众和专业设计师生成高质量高分辨率(HR)图像的重要工具。然而,获得理想结果通常需要多次尝试不同提示词和随机种子生成大量高分辨率图像,带来高昂的计算成本。先生成低分辨率(LR)图像可减轻负担,但如何使低分辨率图像保持与高分辨率版本的感知一致性仍具挑战。本文提出生成高保真低分辨率图像(称为预览图),以在生成最终高分辨率图像前快速识别有潜力的候选方案。针对流匹配模型,我们提出对易零条件,确保低分辨率与高分辨率图像之间的感知一致性,并据此设计无需训练的解决方案:通过选择下采样矩阵并引入对易零引导。大量实验表明,该方法可在保持高分辨率感知一致性的前提下实现最高达33%的计算量降低;当与现有加速技术结合时,可实现最高3倍的速度提升。此外,该方法还可扩展至图像扭曲、平移等操作,展现出良好的泛化能力。
原文摘要 · Abstract (English)
Image generative models have become indispensable tools to yield exquisite high-resolution (HR) images for everyone, ranging from general users to professional designers. However, a desired outcome often requires generating a large number of HR images with different prompts and seeds, resulting in high computational cost for both users and service providers. Generating low-resolution (LR) images first could alleviate computational burden, but it is not straightforward how to generate LR images that are perceptually consistent with their HR counterparts. Here, we consider the task of generating high-fidelity LR images, called Previews, that preserve perceptual similarity of their HR counterparts for an efficient workflow, allowing users to identify promising candidates before generating the final HR image. We propose the commutator-zero condition to ensure the LR-HR perceptual consistency for flow matching models, leading to the proposed training-free solution with downsampling matrix selection and commutator-zero guidance. Extensive experiments show that our method can generate LR images with up to 33\% computation reduction while maintaining HR perceptual consistency. When combined with existing acceleration techniques, our method achieves up to 3$\times$ speedup. Moreover, our formulation can be extended to image manipulations, such as warping and translation, demonstrating its generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。