arXiv:2510.07631cs.CV2025-10NeurIPS被引 17

解决流模型生成中引导失效问题,提升文本图像生成稳定性。

Rectified-CFG++ for Flow Based Models

  • 提出自适应预测-校正框架,结合确定性流与几何感知条件规则。
  • 在多个大模型上实现比标准CFG更优的生成质量与鲁棒性。
  • 适合追求高稳定性和细节精度的文本到图像生成研究者。

分类器自由引导(CFG)是驱动大型扩散模型实现文本条件生成的核心方法,但其直接应用于基于修正流(RF)的模型时会导致严重偏离数据流形,引发视觉伪影、文本对齐错误和脆弱行为。本文提出Rectified-CFG++,一种自适应的预测-校正引导机制,将修正流的确定性高效性与几何感知的条件规则相结合。推理每一步先执行条件化修正流更新,将样本锚定在学习到的传输路径附近,再施加加权条件校正,插值于条件与无条件速度场之间。我们证明该速度场具有边际一致性,其轨迹始终位于数据流形的有界管状邻域内,确保在广泛引导强度下保持稳定。在大规模文本到图像模型(Flux、Stable Diffusion 3/3.5、Lumina)上的大量实验表明,Rectified-CFG++在MS-COCO、LAION-Aesthetic和T2I-CompBench等基准数据集上持续优于标准CFG。

原文摘要 · Abstract (English)

Classifier-free guidance (CFG) is the workhorse for steering large diffusion models toward text-conditioned targets, yet its native application to rectified flow (RF) based models provokes severe off-manifold drift, yielding visual artifacts, text misalignment, and brittle behaviour. We present Rectified-CFG++, an adaptive predictor-corrector guidance that couples the deterministic efficiency of rectified flows with a geometry-aware conditioning rule. Each inference step first executes a conditional RF update that anchors the sample near the learned transport path, then applies a weighted conditional correction that interpolates between conditional and unconditional velocity fields. We prove that the resulting velocity field is marginally consistent and that its trajectories remain within a bounded tubular neighbourhood of the data manifold, ensuring stability across a wide range of guidance strengths. Extensive experiments on large-scale text-to-image models (Flux, Stable Diffusion 3/3.5, Lumina) show that Rectified-CFG++ consistently outperforms standard CFG on benchmark datasets such as MS-COCO, LAION-Aesthetic, and T2I-CompBench. Project page: https://rectified-cfgpp.github.io/

生成模型扩散模型文本生成流模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。