arXiv:2504.06232cs.CV2025-04NeurIPS被引 20

无需训练,通过流对齐引导实现高质量高分辨率图像生成。

HiFlow: Training-free High-Resolution Image Generation with Flow-Aligned Guidance

  • 利用虚拟参考流对齐高低分辨率间的生成轨迹,提升结构一致性。
  • 在多个基准上生成图像质量超越现有方法,细节更清晰、无明显伪影。
  • 适用于各类预训练扩散模型,特别适合追求高分辨率输出的研究者。

文本到图像(T2I)扩散/流模型因强大的视觉生成能力受到广泛关注。然而,高分辨率图像合成面临内容稀缺与复杂度高的挑战。近期方法尝试使用预训练模型进行免训练的高分辨率生成,但常因仅依赖低分辨率采样路径终点而忽略中间状态,导致图像质量差、出现伪影或细节模糊。为此,本文提出HiFlow,一种免训练、模型无关的框架,以释放预训练流模型的分辨率潜力。HiFlow在高分辨率空间构建虚拟参考流,有效捕捉低分辨率流特征,并通过三方面提供指导:初始化对齐以保持低频一致性、方向对齐以保留结构、加速对齐以增强细节保真度。借助这种流对齐引导,HiFlow显著提升T2I模型的高分辨率生成质量,并在多种个性化变体中表现出色。大量实验验证其在主流指标上优于现有最先进方法。

原文摘要 · Abstract (English)

Text-to-image (T2I) diffusion/flow models have drawn considerable attention recently due to their remarkable ability to deliver flexible visual creations. Still, high-resolution image synthesis presents formidable challenges due to the scarcity and complexity of high-resolution content. Recent approaches have investigated training-free strategies to enable high-resolution image synthesis with pre-trained models. However, these techniques often struggle with generating high-quality visuals and tend to exhibit artifacts or low-fidelity details, as they typically rely solely on the endpoint of the low-resolution sampling trajectory while neglecting intermediate states that are critical for preserving structure and synthesizing finer detail. To this end, we present HiFlow, a training-free and model-agnostic framework to unlock the resolution potential of pre-trained flow models. Specifically, HiFlow establishes a virtual reference flow within the high-resolution space that effectively captures the characteristics of low-resolution flow information, offering guidance for high-resolution generation through three key aspects: initialization alignment for low-frequency consistency, direction alignment for structure preservation, and acceleration alignment for detail fidelity. By leveraging such flow-aligned guidance, HiFlow substantially elevates the quality of high-resolution image synthesis of T2I models and demonstrates versatility across their personalized variants. Extensive experiments validate HiFlow's capability in achieving superior high-resolution image quality over state-of-the-art methods.

图像生成扩散模型高分辨率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。