通过实例感知对齐提升真实场景超分辨率细节恢复能力
InstanceRSR: Real-World Super-Resolution via Instance-Aware Representation Alignment
- 联合建模图像与语义分割,实现实例级特征对齐
- 在多个真实数据集上超越现有方法,达成新SOTA
- 适合需要精细细节还原的图像修复与增强任务
基于生成先验的现有真实世界超分辨率(RSR)方法在生成高质量、全局一致的重建结果方面取得了显著进展,但往往难以恢复复杂真实场景中多样物体实例的细粒度细节。这一局限主要源于常用的去噪损失(如MSE)天然偏好全局一致性,而忽视了实例级感知与修复。为此,我们提出InstanceRSR,一种新型RSR框架,通过联合建模语义信息并引入实例级特征对齐。具体地,以低分辨率(LR)图像为全局一致性引导,同时联合建模图像数据与语义分割图,在采样过程中强化语义相关性。此外,设计实例表示学习模块,将扩散隐空间与实例隐空间对齐,实现实例感知特征对齐,并引入尺度对齐机制以增强细粒度感知与细节恢复能力。得益于这些设计,本方法不仅生成逼真的细节,还在实例级别保持语义一致性。在多个真实世界基准上的大量实验表明,InstanceRSR在定量指标和视觉质量上均显著优于现有方法,达到新的最先进水平。
原文摘要 · Abstract (English)
Existing real-world super-resolution (RSR) methods based on generative priors have achieved remarkable progress in producing high-quality and globally consistent reconstructions. However, they often struggle to recover fine-grained details of diverse object instances in complex real-world scenes. This limitation primarily arises because commonly adopted denoising losses (e.g., MSE) inherently favor global consistency while neglecting instance-level perception and restoration. To address this issue, we propose InstanceRSR, a novel RSR framework that jointly models semantic information and introduces instance-level feature alignment. Specifically, we employ low-resolution (LR) images as global consistency guidance while jointly modeling image data and semantic segmentation maps to enforce semantic relevance during sampling. Moreover, we design an instance representation learning module to align the diffusion latent space with the instance latent space, enabling instance-aware feature alignment, and further incorporate a scale alignment mechanism to enhance fine-grained perception and detail recovery. Benefiting from these designs, our approach not only generates photorealistic details but also preserves semantic consistency at the instance level. Extensive experiments on multiple real-world benchmarks demonstrate that InstanceRSR significantly outperforms existing methods in both quantitative metrics and visual quality, achieving new state-of-the-art (SOTA) performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。