让4K图像生成更稳定清晰,适配各种画幅比例。
UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios
- 从数据到模型协同设计,解决4K图像生成中的定位编码与压缩难题。
- 在4096分辨率下多画幅测试中,各项指标超越主流开源模型。
- 适合追求高细节、跨画幅一致性的图像生成研究与应用者。
扩散变换器在约1K分辨率下已表现出色,但将其扩展至支持多样画幅的原生4K生成时,暴露了位置编码、VAE压缩与优化之间的紧密耦合失效问题。单独处理任一因素均无法充分提升质量。为此,我们采用数据-模型协同设计思路,提出基于Flux的DiT模型UltraFlux,其在包含100万张图像的MultiAspect-4K-1M数据集上原生训练,该数据集覆盖可控多画幅,含双语描述及丰富的视觉语言模型和图像质量评估元数据,支持分辨率与画幅感知采样。模型方面,UltraFlux结合:(i) Resonance 2D RoPE与YaRN实现4K下训练窗口、频率与画幅感知的位置编码;(ii) 简单非对抗性VAE后训练方案提升4K重建保真度;(iii) SNR感知的Huber小波目标函数,在时间步与频带间重平衡梯度;(iv) 分阶段审美课程学习策略,将高审美监督集中于由模型先验主导的高噪声步骤。上述组件共同构建出稳定、保留细节的4K DiT,能泛化至宽屏、正方形与竖屏等多种画幅。在Aesthetic-Eval at 4096基准及多画幅4K设置下,UltraFlux在保真度、美学与对齐性指标上持续优于主流开源基线,并通过大模型提示优化后,达到或超越专有模型Seedream 4.0。
原文摘要 · Abstract (English)
Diffusion transformers have recently delivered strong text-to-image generation around 1K resolution, but we show that extending them to native 4K across diverse aspect ratios exposes a tightly coupled failure mode spanning positional encoding, VAE compression, and optimization. Tackling any of these factors in isolation leaves substantial quality on the table. We therefore take a data-model co-design view and introduce UltraFlux, a Flux-based DiT trained natively at 4K on MultiAspect-4K-1M, a 1M-image 4K corpus with controlled multi-AR coverage, bilingual captions, and rich VLM/IQA metadata for resolution- and AR-aware sampling. On the model side, UltraFlux couples (i) Resonance 2D RoPE with YaRN for training-window-, frequency-, and AR-aware positional encoding at 4K; (ii) a simple, non-adversarial VAE post-training scheme that improves 4K reconstruction fidelity; (iii) an SNR-Aware Huber Wavelet objective that rebalances gradients across timesteps and frequency bands; and (iv) a Stage-wise Aesthetic Curriculum Learning strategy that concentrates high-aesthetic supervision on high-noise steps governed by the model prior. Together, these components yield a stable, detail-preserving 4K DiT that generalizes across wide, square, and tall ARs. On the Aesthetic-Eval at 4096 benchmark and multi-AR 4K settings, UltraFlux consistently outperforms strong open-source baselines across fidelity, aesthetic, and alignment metrics, and-with a LLM prompt refiner-matches or surpasses the proprietary Seedream 4.0.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。