arXiv:2505.05209cs.CV2025-05

用扩散Transformer提升盲超分辨率,效果超越传统方法。

EAM: Enhancing Anything with Diffusion Transformers for Blind Super-Resolution

  • 引入新型Ψ-DiT模块,通过低分辨率潜空间控制增强图像修复
  • 在多个数据集上达到最新性能,视觉质量与定量指标均领先
  • 适合关注图像超分辨率与扩散模型应用的研究者

利用预训练文本到图像(T2I)扩散模型指导盲超分辨率(BSR)已成为主流方法。尽管传统T2I模型依赖U-Net架构,但近期研究表明扩散Transformer(DiT)在此领域表现更优。本文提出增强任意内容模型(EAM),一种基于DiT的新型BSR方法,优于以往U-Net基方法。我们设计了新型Ψ-DiT模块,通过低分辨率潜空间作为可分离流注入控制,构建三流架构,有效利用预训练DiT中的先验知识。为充分挖掘T2I模型的先验引导能力并提升其在BSR中的泛化性,提出渐进式掩码图像建模策略,同时降低训练成本。此外,提出基于多模态模型的主体感知提示生成策略,在上下文学习框架中自动识别关键图像区域,生成详细描述,优化T2I扩散先验的使用。实验表明,EAM在多个数据集上均达到最先进水平,显著优于现有方法,在定量指标和视觉质量方面均有提升。

原文摘要 · Abstract (English)

Utilizing pre-trained Text-to-Image (T2I) diffusion models to guide Blind Super-Resolution (BSR) has become a predominant approach in the field. While T2I models have traditionally relied on U-Net architectures, recent advancements have demonstrated that Diffusion Transformers (DiT) achieve significantly higher performance in this domain. In this work, we introduce Enhancing Anything Model (EAM), a novel BSR method that leverages DiT and outperforms previous U-Net-based approaches. We introduce a novel block, $Ψ$-DiT, which effectively guides the DiT to enhance image restoration. This block employs a low-resolution latent as a separable flow injection control, forming a triple-flow architecture that effectively leverages the prior knowledge embedded in the pre-trained DiT. To fully exploit the prior guidance capabilities of T2I models and enhance their generalization in BSR, we introduce a progressive Masked Image Modeling strategy, which also reduces training costs. Additionally, we propose a subject-aware prompt generation strategy that employs a robust multi-modal model in an in-context learning framework. This strategy automatically identifies key image areas, provides detailed descriptions, and optimizes the utilization of T2I diffusion priors. Our experiments demonstrate that EAM achieves state-of-the-art results across multiple datasets, outperforming existing methods in both quantitative metrics and visual quality.

盲超分辨率扩散模型Transformer图像修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。