提出RF-Sampling,让流模型生成图像更贴合提示词且质量更高。
Reflective Flow Sampling Enhancement
- 基于流反演与文本表示线性组合,实现无需训练的推理增强。
- 在多个基准上显著提升生成质量和提示词对齐度,尤其适配FLUX类模型。
- 首次使流模型具备测试时缩放能力,可动态提升生成效果。
文本到图像生成需求推动生成模型快速发展。近期基于流匹配算法(如FLUX)的文本到图像扩散模型取得了显著进展,成为传统扩散模型的有力替代。同时,推理时增强策略被证明可提升生成质量与提示词对齐度。然而,这些方法主要适用于传统扩散模型,在流模型上表现不佳。为此,我们提出反射流采样(RF-Sampling),一种理论严谨、无需训练的推理增强框架,专为流模型(尤其是从CFG引导技术中蒸馏出的变体,如FLUX)设计。不同于启发式方法,我们提供形式化推导,证明RF-Sampling隐式执行文本-图像对齐得分的梯度上升。通过融合文本表示的线性组合与流反演,该方法使模型探索更契合输入提示的噪声空间。大量实验表明,RF-Sampling在多个基准上持续提升生成质量与提示词对齐度。此外,它是首个能在FLUX上表现出一定测试时缩放能力的推理增强方法。
原文摘要 · Abstract (English)
The growing demand for text-to-image generation has led to rapid advances in generative modeling. Recently, text-to-image diffusion models trained with flow matching algorithms, such as FLUX, have achieved remarkable progress and emerged as strong alternatives to conventional diffusion models. At the same time, inference-time enhancement strategies have been shown to improve the generation quality and text-prompt alignment of text-to-image diffusion models. However, these techniques are mainly applicable to conventional diffusion models and usually fail to perform well on flow models. To bridge this gap, we propose Reflective Flow Sampling (RF-Sampling), a theoretically-grounded and training-free inference enhancement framework explicitly designed for flow models, especially for the CFG-distilled variants (i.e., models distilled from CFG guidance techniques), like FLUX. Departing from heuristic interpretations, we provide a formal derivation proving that RF-Sampling implicitly performs gradient ascent on the text-image alignment score. By leveraging a linear combination of textual representations and integrating them with flow inversion, RF-Sampling allows the model to explore noise spaces that are more consistent with the input prompt. Extensive experiments across multiple benchmarks demonstrate that RF-Sampling consistently improves both generation quality and prompt alignment. Moreover, RF-Sampling is also the first inference enhancement method that can exhibit test-time scaling ability to some extent on FLUX.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。