用注意力机制提升图像超分辨率,让细节更真实、结构更完整。
Local-Global Context-Aware and Structure-Preserving Image Super-Resolution
- 引入局部全局上下文感知注意力,保持像素间关系
- 通过像素级分布对齐,提升感知保真度与结构一致性
- 适合追求真实感和结构准确的图像修复任务
扩散模型在图像处理任务中取得显著进展,包括图像超分辨率和感知质量增强。预训练的文本到图像模型(如Stable Diffusion)展现出强大的真实图像内容生成能力,使其在超分辨率任务中极具吸引力。然而,现有方法在面对多样且严重退化的图像时,常出现噪声放大或内容错误生成。为此,我们提出一种情境精确的图像超分辨率框架,通过局部-全局上下文感知注意力有效维持局部与全局像素关系,实现高质量图像生成。此外,我们在像素空间设计了分布与感知对齐的条件机制,捕捉细粒度像素表示,逐步保留并优化结构信息,从局部细节过渡到全局结构。推理时,该方法生成的图像在结构上与原图一致,减少伪影,确保真实细节恢复。多个超分辨率基准测试表明,本方法能生成高保真、感知准确的重建结果。
原文摘要 · Abstract (English)
Diffusion models have recently achieved significant success in various image manipulation tasks, including image super-resolution and perceptual quality enhancement. Pretrained text-to-image models, such as Stable Diffusion, have exhibited strong capabilities in synthesizing realistic image content, which makes them particularly attractive for addressing super-resolution tasks. While some existing approaches leverage these models to achieve state-of-the-art results, they often struggle when applied to diverse and highly degraded images, leading to noise amplification or incorrect content generation. To address these limitations, we propose a contextually precise image super-resolution framework that effectively maintains both local and global pixel relationships through Local-Global Context-Aware Attention, enabling the generation of high-quality images. Furthermore, we propose a distribution- and perceptual-aligned conditioning mechanism in the pixel space to enhance perceptual fidelity. This mechanism captures fine-grained pixel-level representations while progressively preserving and refining structural information, transitioning from local content details to the global structural composition. During inference, our method generates high-quality images that are structurally consistent with the original content, mitigating artifacts and ensuring realistic detail restoration. Extensive experiments on multiple super-resolution benchmarks demonstrate the effectiveness of our approach in producing high-fidelity, perceptually accurate reconstructions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。