提出可控一步扩散模型CODSR,提升图像超分辨率的保真度与细节质量。
Bridging Fidelity-Reality with Controllable One-Step Diffusion for Image Super-Resolution
- 利用未压缩低质输入信息,增强生成过程的保真条件。
- 区域自适应激活生成先验,提升视觉丰富性且不损失局部结构。
- 文本匹配引导策略对齐提示词与语义区域,适合需要精细控制的应用。
基于扩散的一步超分辨率方法虽取得显著进展,但仍受三大限制:(1)因低质量输入压缩编码导致的信息丢失,造成保真度下降;(2)生成先验在区域上缺乏区分性激活;(3)文本提示与其对应语义区域存在错位。为此,本文提出可控一步扩散模型CODSR。首先设计未压缩低质输入引导的特征调制模块,利用原始未压缩信息为扩散过程提供高保真条件;其次提出区域自适应生成先验激活方法,有效增强感知丰富性而不牺牲局部结构保真度;最后采用文本匹配引导策略,充分挖掘文本提示的条件潜力。大量实验表明,CODSR在保持高效一步推理的同时,相较当前最优方法实现了更优的感知质量和具有竞争力的保真度。
原文摘要 · Abstract (English)
Recent diffusion-based one-step methods have shown remarkable progress in the field of image super-resolution, yet they remain constrained by three critical limitations: (1) inferior fidelity performance caused by the information loss from compression encoding of low-quality (LQ) inputs; (2) insufficient region-discriminative activation of generative priors; (3) misalignment between text prompts and their corresponding semantic regions. To address these limitations, we propose CODSR, a controllable one-step diffusion network for image super-resolution. First, we propose an LQ-guided feature modulation module that leverages original uncompressed information from LQ inputs to provide high-fidelity conditioning for the diffusion process. We then develop a region-adaptive generative prior activation method to effectively enhance perceptual richness without sacrificing local structural fidelity. Finally, we employ a text-matching guidance strategy to fully harness the conditioning potential of text prompts. Extensive experiments demonstrate that CODSR achieves superior perceptual quality and competitive fidelity compared with state-of-the-art methods with efficient one-step inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。