arXiv:2607.09351cs.CV2026-07

用可学习提示提升文本增强超分的精度与鲁棒性

Simon-SR: Spatially Adaptive Modulation and Visual Prompt Adaptation for Text-Reinforced Super-Resolution

论文配图:Simon-SR: Spatially Adaptive Modulation and Visual Prompt Adaptation for Text-Reinforced Super-Resolution
图 1 · 摘自论文原文
  • 引入可学习提示实现高效语义挖掘与文本图像融合
  • 在PSNR上最高提升0.50 dB,SSIM提升0.0133,LPIPS降低0.0695
  • 适合需要低标注成本的多模态图像超分场景

单图像超分辨率(SISR)从低分辨率输入重建高质量图像。尽管近期多模态方法提升了感知质量,但对错误先验敏感且需昂贵标注。为此,我们提出Simon-SR,一种利用可学习提示进行高效语义挖掘和鲁棒文本-图像融合的多模态SISR框架。方法结合对比提示学习与提示引导的空间自适应精炼,强化多模态对齐。实验表明,Simon-SR超越现有最优方法,在PSNR上最大提升0.50 dB,SSIM提升0.0133,LPIPS降低0.0695。代码将公开。

原文摘要 · Abstract (English)

Single Image Super-Resolution (SISR) reconstructs high-quality images from low-resolution inputs. While recent multi-modal methods improve perceptual quality, they remain sensitive to erroneous priors and require expensive annotations. To address these issues, we propose Simon-SR, a multi-modal SISR framework leveraging learnable prompts for efficient semantic mining and robust text-image fusion. Our approach combines Contrastive Prompt Learning with Prompt-Guided Spatially Adaptive Refinement to enhance multi-modal alignment. Experiments demonstrate that Simon-SR surpasses state-of-the-art methods, achieving maximum improvements of 0.50 dB in PSNR, 0.0133 in SSIM, and 0.0695 in LPIPS. Code will be released.

图像超分多模态提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。