arXiv:2601.16660eess.IVcs.CV2026-01被引 1

用流图模型实现快速高保真超分辨率,兼顾真实感与效率。

Fast, faithful and photorealistic diffusion-based image super-resolution with enhanced Flow Map models

  • 基于流图模型改进扩散架构,支持单步快速推理。
  • 在x4和x8放大下,保真度与真实感均优于现有方法。
  • 无需分尺度训练,一套模型通用于多种放大倍数。

基于扩散模型的图像超分辨率近期受到广泛关注,得益于大型文本到图像扩散模型的强大表达能力。核心挑战在于重建保真度与真实感之间的权衡。为提升推理效率,许多工作采用针对超分辨率定制的知识蒸馏策略,实现单步扩散方法。然而,这种师生范式受限于信息压缩,可能损失真实纹理与景深等感知线索,即使整体感知质量较高。与此同时,自蒸馏扩散模型(即流图模型)在图像生成任务中表现出色,能在保持标准扩散模型表达力与训练稳定性的同时实现快速推理。本文提出FlowMapSR,一种专为高效推理设计的新型基于扩散的超分辨率框架。在适配流图模型至超分辨率的基础上,引入两项互补增强:(i) 基于分类器自由引导泛化的正负提示引导机制;(ii) 使用低秩适配(LoRA)进行对抗性微调。在考虑的三种流图形式(欧拉、拉格朗日、捷径)中,捷径变体结合上述增强后表现最佳。大量实验表明,FlowMapSR在x4和x8放大下均实现了比现有最先进方法更优的保真度与真实感平衡,同时保持了竞争力的推理速度。值得注意的是,单一模型适用于两种放大因子,无需特定尺度条件或退化引导机制。

原文摘要 · Abstract (English)

Diffusion-based image super-resolution (SR) has recently attracted significant attention by leveraging the expressive power of large pre-trained text-to-image diffusion models (DMs). A central practical challenge is resolving the trade-off between reconstruction faithfulness and photorealism. To address inference efficiency, many recent works have explored knowledge distillation strategies specifically tailored to SR, enabling one-step diffusion-based approaches. However, these teacher-student formulations are inherently constrained by information compression, which can degrade perceptual cues such as lifelike textures and depth of field, even with high overall perceptual quality. In parallel, self-distillation DMs, known as Flow Map models, have emerged as a promising alternative for image generation tasks, enabling fast inference while preserving the expressivity and training stability of standard DMs. Building on these developments, we propose FlowMapSR, a novel diffusion-based framework for image super-resolution explicitly designed for efficient inference. Beyond adapting Flow Map models to SR, we introduce two complementary enhancements: (i) positive-negative prompting guidance, based on a generalization of classifier free-guidance paradigm to Flow Map models, and (ii) adversarial fine-tuning using Low-Rank Adaptation (LoRA). Among the considered Flow Map formulations (Eulerian, Lagrangian, and Shortcut), we find that the Shortcut variant consistently achieves the best performance when combined with these enhancements. Extensive experiments show that FlowMapSR achieves a better balance between reconstruction faithfulness and photorealism than recent state-of-the-art methods for both x4 and x8 upscaling, while maintaining competitive inference time. Notably, a single model is used for both upscaling factors, without any scale-specific conditioning or degradation-guided mechanisms.

超分辨率扩散模型流图模型高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。