用语义分割提升扩散模型的图像超分辨率,更准更可控。
HoliSDiP: Image Super-Resolution via Holistic Semantics and Diffusion Prior
- 用语义标签替代模糊文本提示,结合分割图提供精准引导。
- 在多个真实场景下显著提升图像质量,有效减少噪声干扰。
- 适合需要高保真细节重建的图像修复与增强任务。
文本到图像的扩散模型已成为真实世界图像超分辨率(Real-ISR)的强大先验。然而,现有方法因文本提示噪声大且缺乏空间信息,可能导致意外结果。本文提出 HoliSDiP 框架,利用语义分割为基于扩散的 Real-ISR 提供精确的文本与空间引导。该方法采用语义标签作为简洁文本提示,并通过分割掩码及提出的 Segmentation-CLIP Map 实现密集语义引导。大量实验表明,HoliSDiP 在多种 Real-ISR 场景下均实现显著性能提升,归因于提示噪声降低与空间控制增强。
原文摘要 · Abstract (English)
Text-to-image diffusion models have emerged as powerful priors for real-world image super-resolution (Real-ISR). However, existing methods may produce unintended results due to noisy text prompts and their lack of spatial information. In this paper, we present HoliSDiP, a framework that leverages semantic segmentation to provide both precise textual and spatial guidance for diffusion-based Real-ISR. Our method employs semantic labels as concise text prompts while introducing dense semantic guidance through segmentation masks and our proposed Segmentation-CLIP Map. Extensive experiments demonstrate that HoliSDiP achieves significant improvement in image quality across various Real-ISR scenarios through reduced prompt noise and enhanced spatial control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。