用加权h变换引导生成,从低质图像高质量还原细节。
Coarse-Guided Visual Generation via Weighted h-Transform Sampling
- 通过h变换修改采样过程的转移概率,实现无训练引导生成。
- 在图像和视频生成任务中均达到高保真度,优于现有方法。
- 适合需要快速部署、无需配对数据的视觉生成场景。
粗粒度引导的视觉生成可从低质量或退化的粗略参考图像合成精细视觉样本,广泛应用于真实场景。尽管基于训练的方法有效,但受限于高昂的训练成本和因成对数据收集导致的泛化能力差。近期无需训练的方法利用预训练扩散模型,在采样过程中引入引导,但通常需已知前向(精细到粗略)变换算子(如双三次下采样),或难以平衡引导强度与合成质量。为此,我们提出一种新方法,采用h变换这一约束随机过程的工具,通过在每个采样时间步添加漂移函数,修改原始微分方程的转移概率,近似引导生成趋向理想精细样本。为应对不可避免的近似误差,引入噪声级别感知调度机制,随误差增大逐步降低该项权重,兼顾引导准确性与生成质量。在多种图像与视频生成任务上进行的大量实验验证了该方法的有效性与泛化能力。
原文摘要 · Abstract (English)
Coarse-guided visual generation, which synthesizes fine visual samples from degraded or low-fidelity coarse references, is essential for various real-world applications. While training-based approaches are effective, they are inherently limited by high training costs and restricted generalization due to paired data collection. Accordingly, recent training-free works propose to leverage pretrained diffusion models and incorporate guidance during the sampling process. However, these training-free methods either require knowing the forward (fine-to-coarse) transformation operator, e.g., bicubic downsampling, or are difficult to balance between guidance and synthetic quality. To address these challenges, we propose a novel guided method by using the h-transform, a tool that can constrain stochastic processes (e.g., sampling process) under desired conditions. Specifically, we modify the transition probability at each sampling timestep by adding to the original differential equation with a drift function, which approximately steers the generation toward the ideal fine sample. To address unavoidable approximation errors, we introduce a noise-level-aware schedule that gradually de-weights the term as the error increases, ensuring both guidance adherence and high-quality synthesis. Extensive experiments across diverse image and video generation tasks demonstrate the effectiveness and generalization of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。