通过定制语义与多模态引导,提升真实图像超分辨率的细节与结构一致性。
MegaSR: Mining Customized Semantics and Expressive Guidance for Real-World Image Super-Resolution
- 引入定制化语义模块与多模态信号融合,增强跨层级语义对齐。
- 在Real-ISR数据集上,感知质量与保真度均达当前最优水平。
- 适合关注真实场景图像重建的视觉算法研究者使用。
文本到图像(T2I)模型因其丰富的多模态学习隐式知识,推动了真实世界图像超分辨率(Real-ISR)的发展。然而,现有基于T2I的方法存在三类关键问题:局部区域重建错误的细粒度细节缺失、U-Net块间语义不一致导致的干扰性语义解释,以及边缘模糊引起的结构退化。为此,本文提出MegaSR,通过引入细粒度定制语义与表达性多模态引导,实现语义丰富且结构一致的重建。具体地,设计定制语义模块(CSM),从图像模态补充细粒度语义,并调节多层级知识的语义融合以实现不同U-Net块的个性化适配;同时,通过成对比较识别表达性强的多模态信号,提出多模态信号融合模块(MSFM)进行聚合,提升结构一致性。在真实与合成数据集上的大量实验表明,该方法不仅在质量驱动指标上达到领先水平,同时在保真度指标上保持竞争力,平衡了感知真实感与内容忠实性。
原文摘要 · Abstract (English)
Text-to-image (T2I) models have ushered in a new era of real-world image super-resolution (Real-ISR) due to their rich internal implicit knowledge for multimodal learning. Although bringing high-level semantic priors and dense pixel guidance have led to advances in reconstruction, we identified several critical phenomena by analyzing the behavior of existing T2I-based Real-ISR methods: (1) Fine detail deficiency, which ultimately leads to incorrect reconstruction in local regions. (2) Block-wise semantic inconsistency, which results in distracted semantic interpretations across U-Net blocks. (3) Edge ambiguity, which causes noticeable structural degradation. Building upon these observations, we first introduce MegaSR, which enhances the T2I-based Real-ISR models with fine-grained customized semantics and expressive guidance to unlock semantically rich and structurally consistent reconstruction. Then, we propose the Customized Semantics Module (CSM) to supplement fine-grained semantics from the image modality and regulate the semantic fusion between multi-level knowledge to realize customization for different U-Net blocks. Besides the semantic adaptation, we identify expressive multimodal signals through pair-wise comparisons and introduce the Multimodal Signal Fusion Module (MSFM) to aggregate them for structurally consistent reconstruction. Extensive experiments on real-world and synthetic datasets demonstrate the superiority of the method. Notably, it not only achieves state-of-the-art performance on quality-driven metrics but also remains competitive on fidelity-focused metrics, striking a balance between perceptual realism and faithful content reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。