用强化学习动态优化扩散模型采样,提升少步生成质量
Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning

- 用轻量强化学习自动规划采样路径,聚焦关键阶段
- 在3步生成下保持高细节,减少结构失真
- 适合需要快速高质量生成的图像生成场景
一致性蒸馏显著加速了扩散模型推理,但其采样动态仍缺乏深入研究。我们发现:尽管对数正态采样先验在标准迭代生成中表现良好,一致性蒸馏却呈现出不同难度分布(如U形),瓶颈集中在边界阶段而非中间步骤。为应对静态采样在动态学习需求下的局限性,我们提出曲率自适应一致性流匹配(CACFM)。通过将蒸馏建模为动态决策过程,CACFM利用轻量级强化学习代理探测概率流ODE轨迹,构建以效率为导向的课程,优先处理关键区域而无需人工调度。结合流适应型DMD与对抗性一致性目标,该基于强化学习的调度器在FLUX和SDXL等大规模模型上取得当前最优效果,在极端少步条件下缓解结构畸变并保留高频细节。
原文摘要 · Abstract (English)
Consistency distillation has significantly accelerated diffusion-model inference, but its sampling dynamics remain underexplored. We reveal an asymmetry: although Logit-Normal sampling priors work well for standard iterative generation, consistency distillation exhibits a different difficulty profile (e.g., U-shaped), with optimization bottlenecks concentrated at the boundary stages rather than intermediate steps. To address the limitations of static sampling under evolving learning demands, we propose Curvature-Adaptive Consistency Flow Matching (CACFM). By formulating distillation as a dynamic decision process, CACFM uses a lightweight reinforcement learning agent to probe Probability Flow ODE trajectories and construct an efficiency-oriented curriculum that prioritizes critical regions without manual scheduling. Combined with Flow-adapted DMD and adversarial consistency objectives, our RL-based scheduler achieves state-of-the-art results on large-scale models such as FLUX and SDXL, mitigating structural deformities and preserving high-frequency details in extreme few-step regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。