arXiv:2603.14112cs.CV2026-03

用空间语义引导扩散模型,让超分辨率图像更真实且不失真。

Revisiting the Perception-Distortion Trade-off with Spatial-Semantic Guided Super-Resolution

  • 结合空间位置和语义提示,控制生成细节不偏离原图。
  • 在多个数据集上实现感知质量与失真率的更好平衡。
  • 适合需要高保真又逼真细节的图像修复任务。

图像超分辨率旨在重建既具有高感知质量又低失真的高分辨率图像,但受制于感知-失真权衡的固有瓶颈。基于GAN的方法虽降低失真,仍难以生成真实细粒度纹理;而基于扩散模型的方法虽能合成丰富细节,却常偏离输入,产生结构幻觉并降低保真度。如何在不牺牲保真度的前提下利用扩散模型的强大生成先验,成为关键挑战。为此,本文提出SpaSemSR,一种空间-语义引导的扩散框架,包含两种互补引导机制:其一,空间-文本联合引导将对象级空间线索与语义提示融合,对齐视觉与文本结构以减少失真;其二,多编码器设计与语义退化约束的语义增强视觉引导,统一多模态语义先验,在严重退化条件下提升感知真实性。这两种引导通过空间-语义注意力自适应融合至扩散过程,有效抑制失真与幻觉,同时保留扩散模型优势。在多个基准测试上的大量实验表明,SpaSemSR实现了更优的感知-失真平衡,生成兼具真实感与保真度的恢复结果。

原文摘要 · Abstract (English)

Image super-resolution (SR) aims to reconstruct high resolution images with both high perceptual quality and low distortion, but is fundamentally limited by the perception-distortion trade-off. GAN-based SR methods reduce distortion but still struggle with realistic fine-grained textures, whereas diffusion-based approaches synthesize rich details but often deviate from the input, hallucinating structures and degrading fidelity. This tension raises a key challenge: how to exploit the powerful generative priors of diffusion models without sacrificing fidelity. To address this, we propose SpaSemSR, a spatial-semantic guided diffusion framework with two complementary guidances. First, spatial-grounded textual guidance integrates object-level spatial cues with semantic prompts, aligning textual and visual structures to reduce distortion. Second, semantic-enhanced visual guidance with a multi-encoder design and semantic degradation constraints unifies multimodal semantic priors, improving perceptual realism under severe degradations. These complementary guidances are adaptively fused into the diffusion process via spatial-semantic attention, suppressing distortion and hallucination while retaining the strengths of diffusion models. Extensive experiments on multiple benchmarks show that SpaSemSR achieves a superior perception-distortion balance, producing both realistic and faithful restorations.

超分辨率扩散模型语义引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。