arXiv:2603.07430cs.CV2026-03中稿 · CVPR被引 1

分离语义先验提升扩散图像超分的可控性与细节还原能力

Disentangled Textual Priors for Diffusion-based Image Super-Resolution

  • 将语义先验按空间层级和频率维度解耦,分别建模全局结构与局部纹理
  • 在9.5万张图像上构建新数据集,支持细粒度语义引导生成
  • 采用多分支无分类器指导策略,减少幻觉并保持语义一致性

图像超分辨率(SR)旨在从退化的低分辨率输入重建高分辨率图像。尽管基于扩散模型的SR方法具备强大生成能力,但其性能高度依赖于语义先验的组织方式。现有方法常使用纠缠或粗粒度的先验,将全局布局与局部细节混杂,或混淆结构与纹理信息,从而限制了语义可控性和可解释性。本文提出DTPSR,一种新型扩散式超分框架,通过两个互补维度——空间层次(全局与局部)和频率语义(低频与高频)——引入解耦文本先验。通过显式分离这些先验,DTPSR能同时捕捉场景级结构与对象级细节,并获得频率感知的语义引导。对应的嵌入通过专用交叉注意力模块注入,形成反映视觉内容语义粒度的渐进生成流程,从全局布局到细粒度纹理逐步生成。为支持该范式,我们构建了包含约9.5万对图像-文本的DisText-SR大规模数据集,其中描述被精心解耦为全局、低频与高频三类。为进一步增强可控性与一致性,采用频率感知的负提示的多分支无分类器指导策略,以抑制幻觉和语义漂移。在合成与真实世界基准上的大量实验表明,DTPSR在感知质量、保真度和跨多种退化场景的泛化能力方面均表现优异。

原文摘要 · Abstract (English)

Image Super-Resolution (SR) aims to reconstruct high-resolution images from degraded low-resolution inputs. While diffusion-based SR methods offer powerful generative capabilities, their performance heavily depends on how semantic priors are structured and integrated into the generation process. Existing approaches often rely on entangled or coarse-grained priors that mix global layout with local details, or conflate structural and textural cues, thereby limiting semantic controllability and interpretability. In this work, we propose DTPSR, a novel diffusion-based SR framework that introduces disentangled textual priors along two complementary dimensions: spatial hierarchy (global vs. local) and frequency semantics (low- vs. high-frequency). By explicitly separating these priors, DTPSR enables the model to simultaneously capture scene-level structure and object-specific details with frequency-aware semantic guidance. The corresponding embeddings are injected via specialized cross-attention modules, forming a progressive generation pipeline that reflects the semantic granularity of visual content, from global layout to fine-grained textures. To support this paradigm, we construct DisText-SR, a large-scale dataset containing approximately 95,000 image-text pairs with carefully disentangled global, low-frequency, and high-frequency descriptions. To further enhance controllability and consistency, we adopt a multi-branch classifier-free guidance strategy with frequency-aware negative prompts to suppress hallucinations and semantic drift. Extensive experiments on synthetic and real-world benchmarks show that DTPSR achieves high perceptual quality, competitive fidelity, and strong generalization across diverse degradation scenarios.

图像超分扩散模型解耦先验语义控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。