arXiv:2412.02960cs.CV2024-12被引 8

用语义分割指导扩散模型,让超分辨率更准地还原物体

Semantic Segmentation Prior for Diffusion-Based Real-World Super-Resolution

  • 引入语义分割作为扩散模型的额外控制条件
  • 在多个数据集上提升语义结构保留能力,生成更真实图像
  • 适合关注图像细节与语义一致性的视觉重建研究者

真实世界图像超分辨率(Real-ISR)通过利用大规模文本到图像模型,实现了从识别性文本提示中恢复逼真图像的显著进展。然而,这些方法有时无法识别关键物体,导致相关区域语义恢复不准确。此外,同一区域可能对多个提示产生强响应,引发语义歧义。为此,本文提出将语义分割作为扩散模型的附加控制条件。相比文本提示,语义分割通过为每个像素分配类别标签,能更全面感知图像中的显著物体,并通过显式分配物体到空间区域来缓解语义歧义。实际中,受超分辨率与分割相互促进的启发,我们提出SegSR,采用双扩散框架实现超分辨率与分割扩散模型间的交互。具体设计了双模态桥模块,在反向扩散过程中实现信息动态传递,达成双向增益。大量实验表明,SegSR不仅能生成更真实的图像,还能更有效地保持语义结构。

原文摘要 · Abstract (English)

Real-world image super-resolution (Real-ISR) has achieved a remarkable leap by leveraging large-scale text-to-image models, enabling realistic image restoration from given recognition textual prompts. However, these methods sometimes fail to recognize some salient objects, resulting in inaccurate semantic restoration in these regions. Additionally, the same region may have a strong response to more than one prompt and it will lead to semantic ambiguity for image super-resolution. To alleviate the above two issues, in this paper, we propose to consider semantic segmentation as an additional control condition into diffusion-based image super-resolution. Compared to textual prompt conditions, semantic segmentation enables a more comprehensive perception of salient objects within an image by assigning class labels to each pixel. It also mitigates the risks of semantic ambiguities by explicitly allocating objects to their respective spatial regions. In practice, inspired by the fact that image super-resolution and segmentation can benefit each other, we propose SegSR which introduces a dual-diffusion framework to facilitate interaction between the image super-resolution and segmentation diffusion models. Specifically, we develop a Dual-Modality Bridge module to enable updated information flow between these two diffusion models, achieving mutual benefit during the reverse diffusion process. Extensive experiments show that SegSR can generate realistic images while preserving semantic structures more effectively.

超分辨率扩散模型语义分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。