arXiv:2508.16158cs.CV2025-08被引 3

用区域注意力提升图像超分辨率中的细节清晰度

RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution

  • 通过区域注意力机制精准引导文本与图像信息融合
  • 在多物体场景下生成更真实、一致的视觉细节
  • 适合需要高精度细节还原的图像修复与增强任务

大型视觉语言模型(VLMs)丰富的文本信息与预训练文本到图像(T2I)扩散模型强大的生成先验,已在单图像超分辨率(SISR)中取得显著成果。然而,现有方法在生成多个物体场景下的清晰区域细节方面仍面临挑战,主要源于缺乏细粒度区域描述及模型对复杂提示的捕捉能力不足。为此,我们提出区域注意力引导超分辨率(RAGSR)方法,显式提取局部细粒度信息,并通过新型区域注意力机制有效编码,实现更佳的细节表现与整体视觉一致性。具体而言,RAGSR定位图像中物体区域并为每区域生成细粒度描述,形成区域-文本对作为T2I模型的文本先验。引入区域引导注意力机制,确保每个区域-文本对在注意力过程中被恰当考虑,同时避免无关区域-文本对之间的干扰。该机制使文本与图像信息融合更具控制力,有效克服传统SISR技术的局限。在基准数据集上的实验表明,本方法在生成感知真实的视觉细节的同时,保持了上下文一致性,优于现有方法。

原文摘要 · Abstract (English)

The rich textual information of large vision-language models (VLMs) combined with the powerful generative prior of pre-trained text-to-image (T2I) diffusion models has achieved impressive performance in single-image super-resolution (SISR). However, existing methods still face significant challenges in generating clear and accurate regional details, particularly in scenarios involving multiple objects. This challenge primarily stems from a lack of fine-grained regional descriptions and the models' insufficient ability to capture complex prompts. To address these limitations, we propose a Regional Attention Guided Super-Resolution (RAGSR) method that explicitly extracts localized fine-grained information and effectively encodes it through a novel regional attention mechanism, enabling both enhanced detail and overall visually coherent SR results. Specifically, RAGSR localizes object regions in an image and assigns fine-grained caption to each region, which are formatted as region-text pairs as textual priors for T2I models. A regional guided attention is then leveraged to ensure that each region-text pair is properly considered in the attention process while preventing unwanted interactions between unrelated region-text pairs. By leveraging this attention mechanism, our approach offers finer control over the integration of text and image information, thereby effectively overcoming limitations faced by traditional SISR techniques. Experimental results on benchmark datasets demonstrate that our approach exhibits superior performance in generating perceptually authentic visual details while maintaining contextual consistency compared to existing approaches.

图像超分扩散模型注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。