通过边缘注意力机制提升图像超分辨率生成质量
EatGAN: An Edge-Attention Guided Generative Adversarial Network for Single Image Super-Resolution
- 引入边缘注意力机制,显式隐式结合边缘先验指导生成
- 在Manga 109上达40.87 dB和0.073 LPIPS,性能领先
- 适合关注细节重建与模型可解释性的图像处理研究者
单图像超分辨率(SISR)是图像处理中的关键任务,旨在提升成像系统的分辨率。近年来,深度学习使SISR取得显著进展,其中基于GAN的模型因在感知质量上的优异表现脱颖而出。然而,这类模型在重建真实高频细节和稳定训练方面仍面临挑战。为此,本文提出首个同时显式与隐式利用边缘先验的GAN-based SISR模型——EatGAN:(i) 提出归一化边缘注意力(NEA)机制,基于通道仿射与空间门控将边缘先验转化为轻量、可学习的调制参数,并多次注入融合至生成器;(ii) 设计边缘引导的混合残差块,逐步强化多尺度结构一致性;(iii) 采用包含像素、感知、边缘梯度与对抗项的复合生成目标。实验表明,该模型在畸变导向与感知导向基准上均达到一致的最先进水平。尤其在Manga 109数据集上,分别取得40.87 dB和0.073(LPIPS)的指标,证明将图像先验从被动引导重构为可控调制单元,可为可信、高保真超分辨率提供可行路径。
原文摘要 · Abstract (English)
Single-image super-resolution (SISR) is an important task in image processing, aiming to enhance the resolution of imaging systems. Recently, SISR has made a significant leap and achieved promising results with deep learning. GAN-based models stand out among all the deep learning models because of their excellent performance in perceiving quality. However, it is rather difficult for them to reconstruct realistic high-frequency details and achieve stable training. To solve these issues, we introduce an Edge-Attention guided Generative Adversarial Network (EatGAN), the first GAN-based SISR model that simultaneously leverages edge priors both explicitly and implicitly inside the generator, which (i) proposes a Normalized Edge Attention (NEA) mechanism based on channel-affine and spatial gating that transforms edge prior into lightweight, learnable modulation parameters and injects and fuses them multiple times in a (ii) edge-guided hybrid residual block, which progressively enforces structural consistency across scales; and (iii) a composite generator objective combining pixel, perceptual, edge-gradient, and adversarial terms. Experiments show consistent state-of-the-art across distortion-oriented benchmarks and perception oriented benchmarks. Notably, our model achieves 40.87 dB and 0.073 (LPIPS) on Manga 109, which indicates that reframing image priors from passive guidance into a controllable modulation primitive for generators can chart a practical path toward trustworthy, high-fidelity Super-Resolution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。