arXiv:2604.26503cs.CV2026-04被引 2

提出空间自适应多引导机制,解决扩散模型生成细节与失真矛盾。

Delta Score Matters! Spatial Adaptive Multi Guidance in Diffusion Models

论文配图:Delta Score Matters! Spatial Adaptive Multi Guidance in Diffusion Models
图 1 · 摘自论文原文
  • 基于微分几何分析,发现标准引导存在全局线性偏差
  • 动态调整不同区域的引导强度,高能区保守、低能区激进
  • 无需训练且零计算成本,显著提升图像视频质量

扩散模型在生成复杂静态与动态视觉内容方面取得突破性进展,主要得益于无分类器引导(CFG)。然而,标准CFG依赖全局统一标量,导致生成内容难以兼顾语义细节与结构完整性:低引导值无法注入精细语义,高引导值则引发结构退化、色彩过饱和及视频时序不一致。本文通过微分几何视角揭示该问题本质:CFG实质上是沿数据流形的切向线性外推,而真实数据流形高度弯曲,导致均匀线性步长产生严重法向偏离。为确保生成轨迹安全,我们推导出空间自适应引导的理论上限。基于此,提出无需训练、几乎零开销的时空自适应多引导(SAMG)采样算法。SAMG动态计算每点条件引导能量,对高能边界区域施加保守最小引导以保护微纹理,对低能区域采用激进最大引导以增强语义注入。在多种图像(SD 1.5、SDXL、SD3.5 Medium)与视频架构(CogVideoX、ModelScope)上的实验表明,SAMG有效解决细节-伪影困境,在不增加计算负担的前提下,显著提升语义对齐度、结构完整性和时序平滑性。

原文摘要 · Abstract (English)

Diffusion models have achieved remarkable success in synthesizing complex static and temporal visuals, a breakthrough largely driven by Classifier-Free Guidance (CFG). However, despite its pivotal role in aligning generated content with textual prompts, standard CFG relies on a globally uniform scalar. This homogeneous amplification traps models in a well-documented "detail-artifact dilemma": low guidance scales fail to inject intricate semantics, while high scales inevitably cause structural degradation, color over-saturation, and temporal inconsistencies in videos. In this paper, we expose the physical root of this flaw through the lens of differential geometry. By analyzing Tweedie's Formula, we reveal that CFG intrinsically performs a tangential linear extrapolation. Because the natural data manifold is highly curved, this uniform linear step introduces a severe orthogonal deviation. To keep the generation trajectory safely bounded, we formulate a theoretical upper bound for spatial and adaptive guidance. Based on these geometric insights, we propose Spatial Adaptive Multi Guidance (SAMG), a training-free and virtually zero-cost sampling algorithm. SAMG dynamically computes point-wise conditional guidance energy, applying a conservative minimum scale to high-energy boundary regions to preserve delicate micro-textures, while deploying an aggressive maximum scale in low-energy regions to maximize semantic injection. Extensive experiments across diverse image (SD 1.5, SDXL, SD3.5 Medium) and video (CogVideoX, ModelScope) architectures demonstrate that SAMG effectively resolves the detail-artifact dilemma, achieving superior semantic alignment, structural integrity, and temporal smoothness without any computational overhead.

扩散模型生成质量引导机制视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。