arXiv:2503.14503cs.CVcs.AI2025-03CVPR被引 19

用多模态信息提升图像超分辨率,效果更真实

The Power of Context: How Multimodality Improves Image Super-Resolution

  • 在扩散模型中融合深度、分割、边缘和文本等多模态信息
  • 相比现有方法,视觉质量和细节还原度显著提升
  • 可灵活控制输出风格,适合图像修复与生成任务

单图像超分辨率(SISR)因难以从低分辨率输入恢复精细细节并保持感知质量而面临挑战。现有方法通常依赖有限的图像先验,导致效果不佳。本文提出一种新方法,利用深度、分割、边缘和文本提示等多模态上下文信息,在扩散模型框架内学习强大的生成先验。设计了一种灵活的网络架构,可无须修改扩散过程即融合任意数量的模态输入。关键在于通过其他模态的空间信息引导文本条件,有效缓解文本引发的幻觉问题。各模态的引导强度可独立控制,实现风格调节:如通过深度增强虚化效果,或通过分割调整物体突出程度。大量实验表明,本模型超越当前先进生成式SISR方法,显著提升视觉质量与保真度。

原文摘要 · Abstract (English)

Single-image super-resolution (SISR) remains challenging due to the inherent difficulty of recovering fine-grained details and preserving perceptual quality from low-resolution inputs. Existing methods often rely on limited image priors, leading to suboptimal results. We propose a novel approach that leverages the rich contextual information available in multiple modalities -- including depth, segmentation, edges, and text prompts -- to learn a powerful generative prior for SISR within a diffusion model framework. We introduce a flexible network architecture that effectively fuses multimodal information, accommodating an arbitrary number of input modalities without requiring significant modifications to the diffusion process. Crucially, we mitigate hallucinations, often introduced by text prompts, by using spatial information from other modalities to guide regional text-based conditioning. Each modality's guidance strength can also be controlled independently, allowing steering outputs toward different directions, such as increasing bokeh through depth or adjusting object prominence via segmentation. Extensive experiments demonstrate that our model surpasses state-of-the-art generative SISR methods, achieving superior visual quality and fidelity. See project page at https://mmsr.kfmei.com/.

图像超分多模态扩散模型生成先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。