arXiv:2510.22534cs.CV2025-10NeurIPS被引 3

通过空间聚焦注意力提升图像超分语义准确性

SRSR: Enhancing Semantic Accuracy in Real-World Image Super-Resolution with Spatially Re-Focused Text-Conditioning

  • 引入空间聚焦交叉注意力,用视觉掩码引导文本条件
  • 在真实世界数据集上超越7个先进基线模型
  • 适合需要高语义保真度的图像修复与超分任务

基于扩散模型的图像超分辨率方法常因文本条件不准确或不完整,以及交叉注意力易关注无关像素,导致语义错位和幻觉细节。为此,本文提出一种即插即用的空间重聚焦超分辨率(SRSR)框架,包含两个核心组件:首先引入空间聚焦交叉注意力(SRCA),在推理时利用视觉引导的分割掩码精炼文本条件;其次提出空间目标无分类器引导(STCFG),选择性屏蔽未定位像素上的文本影响,防止幻觉生成。在合成与真实世界数据集上的大量实验表明,SRSR在所有数据集上均优于7个先进基线,在两个真实世界基准上于感知质量指标(LPIPS、DISTS)显著领先,证明其在实现高语义保真度与感知质量方面的有效性。

原文摘要 · Abstract (English)

Existing diffusion-based super-resolution approaches often exhibit semantic ambiguities due to inaccuracies and incompleteness in their text conditioning, coupled with the inherent tendency for cross-attention to divert towards irrelevant pixels. These limitations can lead to semantic misalignment and hallucinated details in the generated high-resolution outputs. To address these, we propose a novel, plug-and-play spatially re-focused super-resolution (SRSR) framework that consists of two core components: first, we introduce Spatially Re-focused Cross-Attention (SRCA), which refines text conditioning at inference time by applying visually-grounded segmentation masks to guide cross-attention. Second, we introduce a Spatially Targeted Classifier-Free Guidance (STCFG) mechanism that selectively bypasses text influences on ungrounded pixels to prevent hallucinations. Extensive experiments on both synthetic and real-world datasets demonstrate that SRSR consistently outperforms seven state-of-the-art baselines in standard fidelity metrics (PSNR and SSIM) across all datasets, and in perceptual quality measures (LPIPS and DISTS) on two real-world benchmarks, underscoring its effectiveness in achieving both high semantic fidelity and perceptual quality in super-resolution.

图像超分扩散模型语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。