用扩散Transformer提升盲超分辨率,效果超越传统方法。
EAM: Enhancing Anything with Diffusion Transformers for Blind Super-Resolution
- 引入新型Ψ-DiT模块,通过低分辨率潜空间控制增强图像修复
- 在多个数据集上达到最新性能,视觉质量与定量指标均领先
- 适合关注图像超分辨率与扩散模型应用的研究者
利用预训练文本到图像(T2I)扩散模型指导盲超分辨率(BSR)已成为主流方法。尽管传统T2I模型依赖U-Net架构,但近期研究表明扩散Transformer(DiT)在此领域表现更优。本文提出增强任意内容模型(EAM),一种基于DiT的新型BSR方法,优于以往U-Net基方法。我们设计了新型Ψ-DiT模块,通过低分辨率潜空间作为可分离流注入控制,构建三流架构,有效利用预训练DiT中的先验知识。为充分挖掘T2I模型的先验引导能力并提升其在BSR中的泛化性,提出渐进式掩码图像建模策略,同时降低训练成本。此外,提出基于多模态模型的主体感知提示生成策略,在上下文学习框架中自动识别关键图像区域,生成详细描述,优化T2I扩散先验的使用。实验表明,EAM在多个数据集上均达到最先进水平,显著优于现有方法,在定量指标和视觉质量方面均有提升。
原文摘要 · Abstract (English)
Utilizing pre-trained Text-to-Image (T2I) diffusion models to guide Blind Super-Resolution (BSR) has become a predominant approach in the field. While T2I models have traditionally relied on U-Net architectures, recent advancements have demonstrated that Diffusion Transformers (DiT) achieve significantly higher performance in this domain. In this work, we introduce Enhancing Anything Model (EAM), a novel BSR method that leverages DiT and outperforms previous U-Net-based approaches. We introduce a novel block, $Ψ$-DiT, which effectively guides the DiT to enhance image restoration. This block employs a low-resolution latent as a separable flow injection control, forming a triple-flow architecture that effectively leverages the prior knowledge embedded in the pre-trained DiT. To fully exploit the prior guidance capabilities of T2I models and enhance their generalization in BSR, we introduce a progressive Masked Image Modeling strategy, which also reduces training costs. Additionally, we propose a subject-aware prompt generation strategy that employs a robust multi-modal model in an in-context learning framework. This strategy automatically identifies key image areas, provides detailed descriptions, and optimizes the utilization of T2I diffusion priors. Our experiments demonstrate that EAM achieves state-of-the-art results across multiple datasets, outperforming existing methods in both quantitative metrics and visual quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。