改进Stable Diffusion生成细节与清晰度,无需重训练。
CADE 2.5 - ZeResFDG: Frequency-Decoupled, Rescaled and Zero-Projected Guidance for SD/SDXL Latent Diffusion Models
- 通过频段解耦、能量重标和零投影三策略优化引导信号。
- 在中等引导强度下提升图像锐度与提示符合度,减少伪影。
- 适合追求高质量生成且希望避免模型重训练的用户。
我们提出CADE 2.5(Comfy Adaptive Detail Enhancer),一个面向SD/SDXL潜空间扩散模型的采样级引导堆栈。核心模块ZeResFDG整合了三项机制:(i) 频率解耦引导,对引导信号的低频与高频成分分别加权;(ii) 能量重标,使有指导预测的样本幅值与正向分支匹配;(iii) 零投影,移除与无条件方向平行的分量。采用带有迟滞效应的轻量级谱式EMA,在采样过程中根据结构形成状态自动切换保守模式与细节增强模式。在各类SD/SDXL采样器上,ZeResFDG在不进行任何再训练的前提下,显著提升锐度、提示遵循度与伪影控制能力。此外,引入一种免训练的推理时稳定器——QSilk Micrograin Stabilizer(分位数钳位+深度/边缘门控微细节注入),有效提升鲁棒性,并在高分辨率下生成自然的高频微纹理,计算开销极小。附录简要讨论该方法对其他参数化方式(如速度)的兼容性,但本文聚焦于SD/SDXL潜空间扩散模型。
原文摘要 · Abstract (English)
We introduce CADE 2.5 (Comfy Adaptive Detail Enhancer), a sampler-level guidance stack for SD/SDXL latent diffusion models. The central module, ZeResFDG, unifies (i) frequency-decoupled guidance that reweights low- and high-frequency components of the guidance signal, (ii) energy rescaling that matches the per-sample magnitude of the guided prediction to the positive branch, and (iii) zero-projection that removes the component parallel to the unconditional direction. A lightweight spectral EMA with hysteresis switches between a conservative and a detail-seeking mode as structure crystallizes during sampling. Across SD/SDXL samplers, ZeResFDG improves sharpness, prompt adherence, and artifact control at moderate guidance scales without any retraining. In addition, we employ a training-free inference-time stabilizer, QSilk Micrograin Stabilizer (quantile clamp + depth/edge-gated micro-detail injection), which improves robustness and yields natural high-frequency micro-texture at high resolutions with negligible overhead. For completeness we note that the same rule is compatible with alternative parameterizations (e.g., velocity), which we briefly discuss in the Appendix; however, this paper focuses on SD/SDXL latent diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。