在参数和算力限制下,实现超越Real-ESRGAN的高效感知超分。
Efficient Perceptual Image Super Resolution: AIM 2025 Study and Benchmark
- 在500万参数、2000 GFLOPs内优化感知质量。
- 新基准上超越Real-ESRGAN,4K图像复原更真实。
- 适合部署于移动端或边缘设备的高质量超分应用。
本文提出对高效感知超分辨率(EPSR)的全面研究与基准测试。尽管高效峰值信噪比(PSNR)超分已有显著进展,但聚焦感知质量的方法仍普遍效率不足。针对这一空白,本文旨在复现或超越Real-ESRGAN的感知效果,同时满足严格效率约束:模型参数不超过500万,计算量不超过2000 GFLOPs(以960x540输入为准)。所提方法在全新构建的测试集上评估,该数据集包含500张4K分辨率图像,每张经多种退化类型处理,且不提供原始高清版本。此设计模拟真实部署场景,构成多样且具挑战性的基准。最优方法在所有基准数据集上均优于Real-ESRGAN,证明了高效方法在感知域的潜力。本文确立了现代高效感知超分的新基线。
原文摘要 · Abstract (English)
This paper presents a comprehensive study and benchmark on Efficient Perceptual Super-Resolution (EPSR). While significant progress has been made in efficient PSNR-oriented super resolution, approaches focusing on perceptual quality metrics remain relatively inefficient. Motivated by this gap, we aim to replicate or improve the perceptual results of Real-ESRGAN while meeting strict efficiency constraints: a maximum of 5M parameters and 2000 GFLOPs, calculated for an input size of 960x540 pixels. The proposed solutions were evaluated on a novel dataset consisting of 500 test images of 4K resolution, each degraded using multiple degradation types, without providing the original high-quality counterparts. This design aims to reflect realistic deployment conditions and serves as a diverse and challenging benchmark. The top-performing approach manages to outperform Real-ESRGAN across all benchmark datasets, demonstrating the potential of efficient methods in the perceptual domain. This paper establishes the modern baselines for efficient perceptual super resolution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。