arXiv:2606.31363cs.CV2026-06

用真实模糊图像和语言模型,让低分辨率图变清晰。

Language-Assisted Super-Resolution from Real-World Low-Resolution Patches

论文配图:Language-Assisted Super-Resolution from Real-World Low-Resolution Patches
图 1 · 摘自论文原文
  • 从真实图像中提取模糊区域,用语言模型引导重建
  • 在语义空间中对齐高低分辨率图像,提升真实感
  • 适合做真实场景图像增强的研究者和开发者

单图像超分辨率旨在从低分辨率(LR)输入重建高分辨率(HR)图像。传统方法依赖成对的HR-LR数据训练,但现实中难以获取。因此多数方法通过手工降质或相机ISP模拟生成LR图像,这些合成退化无法捕捉真实图像的复杂性,导致实际应用泛化能力差。本文观察到,高质量图像中不同深度区域呈现不同分辨率,远距离区域自然形成真实退化的LR块,而近距离为HR块。由此可从真实图像中提取无配对的LR块。针对此问题,提出LA-SR(Language Assistant for SR),一种新型无配对超分辨率框架。核心思想是将无配对超分辨率重构为语言空间中的任务,利用视觉-语言模型弥合LR-HR差距。LA-SR将图像映射至富含语义与质量信息的空间,并引入两种语言引导损失:语言内容损失以保持语义一致性,语言质量损失以增强感知真实性。该对齐机制使LA-SR能有效处理真实LR输入,生成具有高度真实感的输出,克服了仅基于合成数据训练方法的局限。

原文摘要 · Abstract (English)

Single image super-resolution aims to reconstruct high-resolution (HR) images from low-resolution (LR) inputs. Training SR models typically requires paired HR-LR data, which is difficult to obtain in reality. As a result, most methods synthesize LR images by artificially degrading HR images with handcrafted kernels or camera ISP adjustments. However, these synthetic degradations fail to capture the complexity of real LR images, leading to poor generalization in practice. To address this, we observe that even within a single high-quality image, regions at different depths exhibit varying resolutions, where distant regions act as LR patches and closer ones as HR patches. This allows the extraction of real, degradation-induced LR patches from real images. Since these LR patches lack paired HR counterparts, we propose LA-SR (Language Assistant for SR), a novel framework for unpaired SR. The key idea of LA-SR is to redefine unpaired SR in the language space, using vision-language models to bridge the LR-HR gap. LA-SR projects images into a semantically rich space representing both content and quality, and applies two language-guided losses: linguistic content loss to preserve semantic fidelity, and linguistic quality loss to enhance perceptual realism. With this alignment, LA-SR effectively super-resolves real LR inputs, producing realistic outputs that overcome the limitations of synthetic-data-trained methods.

图像超分辨视觉语言模型真实退化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。