用CLIP提升图像超分跨域能力,小样本快速适应新场景。
CLIP-aware Domain-Adaptive Super-Resolution
- 结合CLIP语义与元学习,实现跨域特征对齐与快速适配。
- ×8时比现有方法提升0.15dB,×16时最高增益达0.30dB。
- 适合需要小样本快速部署的超分辨率实际应用。
本文提出一种新型框架CLIP-aware Domain-Adaptive Super-Resolution(CDASR),解决单图超分辨率中的域泛化难题。通过利用CLIP(对比语言-图像预训练)的语义能力,CDASR在多种域和极端放大倍数下实现前所未有的性能。该方法融合了CLIP引导的特征对齐机制与受元学习启发的少样本适应策略,实现高效知识迁移与快速域适应。一个自定义的域自适应模块通过多阶段变换过程(包括CLIP特征处理、空间特征生成与特征融合),将语义信息有效融入超分流程。此外,CDASR采用包含像素级重建、感知相似性和语义一致性在内的多组件损失函数。大量基准测试表明其优越性:在Urban100数据集上,×8放大时相比现有方法提升0.15dB PSNR,×16放大时最高提升达0.30dB。
原文摘要 · Abstract (English)
This work introduces CLIP-aware Domain-Adaptive Super-Resolution (CDASR), a novel framework that addresses the critical challenge of domain generalization in single image super-resolution. By leveraging the semantic capabilities of CLIP (Contrastive Language-Image Pre-training), CDASR achieves unprecedented performance across diverse domains and extreme scaling factors. The proposed method integrates CLIP-guided feature alignment mechanism with a meta-learning inspired few-shot adaptation strategy, enabling efficient knowledge transfer and rapid adaptation to target domains. A custom domain-adaptive module processes CLIP features alongside super-resolution features through a multi-stage transformation process, including CLIP feature processing, spatial feature generation, and feature fusion. This intricate process ensures effective incorporation of semantic information into the super-resolution pipeline. Additionally, CDASR employs a multi-component loss function that combines pixel-wise reconstruction, perceptual similarity, and semantic consistency. Extensive experiments on benchmark datasets demonstrate CDASR's superiority, particularly in challenging scenarios. On the Urban100 dataset at $\times$8 scaling, CDASR achieves a significant PSNR gain of 0.15dB over existing methods, with even larger improvements of up to 0.30dB observed at $\times$16 scaling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。