无需训练,通过注入噪声条件提升图像生成质量
DICT: Data Injection and Contrastive Trajectory Refinement for Conditional Image Generation with Diffusion Models

- 在去噪早期注入噪声条件信号,自适应提取关键信息
- 通过对比相邻去噪阶段优化生成轨迹,提升图像清晰度
- 通用框架适用于风格迁移、超分等多任务,效果优于现有方法
扩散模型已成为条件图像生成的主流范式,但现有方法多分为两类:任务特异性设计虽能提升性能却限制泛化能力;或采用无训练损失引导,将丰富条件压缩为标量目标并逐步引导,导致信息瓶颈与误差累积。针对跨多样化条件生成任务的统一框架需求,本文提出数据注入与对比轨迹精炼(DICT),一种无需训练的推理方法,可在不引入任务依赖架构的前提下增强条件图像生成。DICT引入数据注入机制,在早期去噪阶段融合噪声扰动的条件信号,通过对这些注入信号的引导去噪,自适应地从原始条件中选择并提炼任务相关特征,有效保留空间细节并确保条件与生成结果精准对齐。此外,DICT在相邻去噪状态间实施对比轨迹精炼,通过成对比较逐步提升样本质量。该设计保持推理简洁,同时在统一扩散框架下实现跨任务迁移能力。在风格迁移、图像超分辨率和图像去模糊等任务上的大量实验表明,DICT在保真度和感知质量上均持续优于代表性任务特异性与损失引导基线方法。
原文摘要 · Abstract (English)
Diffusion models have become a dominant paradigm for conditional image generation, yet existing approaches generally follow two directions: task-specific designs that can improve performance but limit generalization, and training-free loss guidance that compresses rich conditions into scalar objectives and applies stepwise guidance, leading to information bottlenecks and error accumulation along the sampling trajectory. Given the urgent need for an effective unified framework across diverse conditional image generation tasks, we propose Data Injection and Contrastive Trajectory Refinement (DICT), a training-free inference method that enhances conditional image generation without introducing task-dependent architectures. DICT introduces Data Injection, where noise-perturbed conditional signals are integrated into early denoising stages; by performing guided denoising on these injected signals, DICT adaptively selects and distills task-salient information from the raw condition, effectively preserving spatial richness and ensuring precise condition-to-generation alignment. Furthermore, DICT applies Contrastive Trajectory Refinement across adjacent denoising states, enabling pairwise comparisons that progressively improve sample quality. These designs keep inference simple while improving cross-task transfer under a unified diffusion formulation. Extensive experiments on conditional image generation tasks (e.g., style transfer, image super-resolution, and image deblurring) show consistent gains in fidelity and perceptual quality over representative task-specific and loss-guided baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。