QSilk提升扩散模型细节清晰度,抑制异常激活且无需训练
QSilk: Micrograin Stabilization and Adaptive Quantile Clipping for Detail-Friendly Latent Diffusion
- 每样本微钳制+自适应分位数裁剪,动态控制输出范围
- 低步数、超高清下结果更锐利,纹理保留更好
- 即插即用无开销,适合图像生成与渲染场景
我们提出QSilk,一种轻量级、始终启用的潜在扩散稳定层,可在不损失纹理的前提下提升高频保真度并抑制罕见激活峰值。QSilk结合了(i)每样本微钳制,温和限制极端值而不模糊纹理;以及(ii)自适应分位数裁剪(AQClip),根据区域特性动态调整允许值区间。AQClip可基于局部结构统计运行于代理模式,或在注意力熵引导模式(模型置信度)下工作。集成至CADE 2.5渲染管线后,QSilk在低步数和超高清场景下实现更清晰、更锐利的结果,开销极小。无需训练或微调,用户控制极少。我们在SD/SDXL主干网络上均观察到一致的定性改进,并与CFG/Rescale协同良好,可在无伪影前提下实现稍高的引导强度。
原文摘要 · Abstract (English)
We present QSilk, a lightweight, always-on stabilization layer for latent diffusion that improves high-frequency fidelity while suppressing rare activation spikes. QSilk combines (i) a per-sample micro clamp that gently limits extreme values without washing out texture, and (ii) Adaptive Quantile Clip (AQClip), which adapts the allowed value corridor per region. AQClip can operate in a proxy mode using local structure statistics or in an attention entropy guided mode (model confidence). Integrated into the CADE 2.5 rendering pipeline, QSilk yields cleaner, sharper results at low step counts and ultra-high resolutions with negligible overhead. It requires no training or fine-tuning and exposes minimal user controls. We report consistent qualitative improvements across SD/SDXL backbones and show synergy with CFG/Rescale, enabling slightly higher guidance without artifacts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。