arXiv:2504.09887cs.CV2025-04CVPR被引 4

用语义引导扩散模型提升真实用户生成图像的超分辨率效果

Enhanced Semantic Extraction and Guidance for UGC Image Super Resolution

  • 将语义提取模型SAM2融入扩散框架,增强细节生成能力
  • 在LSDIR数据集上分离模拟退化过程,提升对真实退化的建模
  • 在CVPR NTIRE 2025挑战赛中获第二名,适合真实场景图像修复

由于真实用户生成内容(UGC)图像与合成退化之间的差异,传统超分辨率方法难以有效泛化,亟需更鲁棒的方法来建模真实退化。本文提出一种新型UGC图像超分辨率方法,将语义引导引入扩散框架。通过在LSDIR数据集上独立模拟退化过程,并与官方成对训练集结合,缓解了真实与合成数据退化不一致的问题。此外,借助预训练语义提取模型SAM2并微调关键超参数,提升了退化去除与细节生成能力,增强了感知保真度。大量实验表明,该方法优于当前最优方法。同时,该模型在CVPR NTIRE 2025短格式UGC图像超分辨率挑战赛中获得第二名,进一步验证其有效性。代码已开源:https://github.com/Moonsofang/NTIRE-2025-SRlab。

原文摘要 · Abstract (English)

Due to the disparity between real-world degradations in user-generated content(UGC) images and synthetic degradations, traditional super-resolution methods struggle to generalize effectively, necessitating a more robust approach to model real-world distortions. In this paper, we propose a novel approach to UGC image super-resolution by integrating semantic guidance into a diffusion framework. Our method addresses the inconsistency between degradations in wild and synthetic datasets by separately simulating the degradation processes on the LSDIR dataset and combining them with the official paired training set. Furthermore, we enhance degradation removal and detail generation by incorporating a pretrained semantic extraction model (SAM2) and fine-tuning key hyperparameters for improved perceptual fidelity. Extensive experiments demonstrate the superiority of our approach against state-of-the-art methods. Additionally, the proposed model won second place in the CVPR NTIRE 2025 Short-form UGC Image Super-Resolution Challenge, further validating its effectiveness. The code is available at https://github.c10pom/Moonsofang/NTIRE-2025-SRlab.

图像超分辨率扩散模型语义引导UGC图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。