无需调参提升图像生成分辨率,稳定增强细节与清晰度。
I-Max: Maximize the Resolution Potential of Pre-trained Rectified Flow Transformers with Projected Flow
- 提出投影流策略实现高分辨率生成的稳定性
- 在Lumina-Next-2K和Flux.1-dev上实现细节涌现与伪影修正
- 适合追求高效高分辨生成的开发者与研究者
Rectified Flow Transformers(RFTs)具备优异的训练与推理效率,是扩展扩散模型最具潜力的方向。然而,由于数据质量与训练成本限制,生成分辨率的提升进展缓慢。无调参的分辨率外推提供了替代方案,但现有方法常降低生成稳定性,制约实际应用。本文系统回顾现有分辨率外推方法,提出I-Max框架以最大化文本到图像RFT的分辨率潜力。I-Max包含:(i) 一种新型投影流策略,保障外推稳定性;(ii) 高级推理工具包,实现模型知识向更高分辨率的泛化。在Lumina-Next-2K与Flux.1-dev上的实验表明,I-Max不仅能提升外推稳定性,还能实现图像细节涌现与伪影修正,验证了无调参分辨率外推的实际价值。
原文摘要 · Abstract (English)
Rectified Flow Transformers (RFTs) offer superior training and inference efficiency, making them likely the most viable direction for scaling up diffusion models. However, progress in generation resolution has been relatively slow due to data quality and training costs. Tuning-free resolution extrapolation presents an alternative, but current methods often reduce generative stability, limiting practical application. In this paper, we review existing resolution extrapolation methods and introduce the I-Max framework to maximize the resolution potential of Text-to-Image RFTs. I-Max features: (i) a novel Projected Flow strategy for stable extrapolation and (ii) an advanced inference toolkit for generalizing model knowledge to higher resolutions. Experiments with Lumina-Next-2K and Flux.1-dev demonstrate I-Max's ability to enhance stability in resolution extrapolation and show that it can bring image detail emergence and artifact correction, confirming the practical value of tuning-free resolution extrapolation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。