arXiv:2508.09449cs.CV2025-08中稿 · ISCAS 2026被引 1

用自动检索参考图提升图像超分辨率,让手机拍照也能清晰还原细节。

RASR: Retrieval-Augmented Super Resolution for Practical Reference-based Image Restoration

  • 基于语义检索从数据库中自动找参考图,不再依赖人工配对。
  • 在新数据集上比传统方法提升0.38 dB PSNR,纹理更真实。
  • 适合手机拍照、博物馆等现实场景,实用性强。

参考图像超分辨率(RefSR)通过利用高质量参考图像,提升了单图像超分辨率(SISR)的纹理保真度和视觉真实性。然而现有RefSR方法严重依赖人工标注的目标-参考图像对,限制了其在真实场景中的应用。为此,本文提出检索增强型超分辨率(RASR),一种可自动从参考数据库中检索与输入低质图像语义相关的高分辨率图像的新范式,实现可扩展、灵活的现实场景应用,如在动物园或博物馆拍摄的手机照片增强。为推动该方向研究,我们构建了首个RASR专用基准数据集RASR-Flickr30,其提供按类别组织的参考图像库,支持开放世界检索。同时提出RASRNet作为强基线模型,结合语义参考检索器与基于扩散的生成器,通过语义条件增强生成效果。在RASR-Flickr30上的实验表明,RASRNet持续优于SISR基线,实现+0.38 dB PSNR和-0.0131 LPIPS,生成更真实的纹理,验证了检索增强是连接学术研究与实际应用的关键路径。

原文摘要 · Abstract (English)

Reference-based Super Resolution (RefSR) improves upon Single Image Super Resolution (SISR) by leveraging high-quality reference images to enhance texture fidelity and visual realism. However, a critical limitation of existing RefSR approaches is their reliance on manually curated target-reference image pairs, which severely constrains their practicality in real-world scenarios. To overcome this, we introduce Retrieval-Augmented Super Resolution (RASR), a new and practical RefSR paradigm that automatically retrieves semantically relevant high-resolution images from a reference database given only a low-quality input. This enables scalable and flexible RefSR in realistic use cases, such as enhancing mobile photos taken in environments like zoos or museums, where category-specific reference data (e.g., animals, artworks) can be readily collected or pre-curated. To facilitate research in this direction, we construct RASR-Flickr30, the first benchmark dataset designed for RASR. Unlike prior datasets with fixed target-reference pairs, RASR-Flickr30 provides per-category reference databases to support open-world retrieval. We further propose RASRNet, a strong baseline that combines a semantic reference retriever with a diffusion-based RefSR generator. It retrieves relevant references based on semantic similarity and employs a diffusion-based generator enhanced with semantic conditioning. Experiments on RASR-Flickr30 demonstrate that RASRNet consistently improves over SISR baselines, achieving +0.38 dB PSNR and -0.0131 LPIPS, while generating more realistic textures. These findings highlight retrieval augmentation as a promising direction to bridge the gap between academic RefSR research and real-world applicability.

图像超分辨率参考图像检索增强扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。