arXiv:2506.03131cs.CVcs.LG2025-06NeurIPS被引 13

提出原生分辨率图像生成,可任意缩放和比例生成高清图。

Native-Resolution Image Synthesis

  • 用新型Transformer架构直接处理不同分辨率与长宽比的图像
  • 单模型在256x256和512x512上均达顶尖性能
  • 仅用ImageNet训练即可零样本生成1536x1536等新尺寸图像

我们提出原生分辨率图像生成,一种可生成任意分辨率与长宽比图像的新范式。该方法突破传统固定分辨率、正方形图像的局限,通过原生处理可变长度视觉标记,解决传统技术核心难题。为此,我们设计了原生分辨率扩散Transformer(NiT),其去噪过程显式建模不同分辨率与长宽比。不受固定格式限制,NiT从广泛分辨率与长宽比的图像中学习内在视觉分布。值得注意的是,单一NiT模型同时在ImageNet-256x256和512x512基准上达到当前最优表现。令人意外的是,类似先进大语言模型的强零样本能力,仅基于ImageNet训练的NiT展现出卓越零样本泛化性能:成功生成此前未见的高分辨率(如1536x1536)及多样化长宽比(如16:9、3:1、4:3)的高质量图像,如图1所示。这些发现表明,原生分辨率建模具有连接视觉生成与先进大语言模型方法的巨大潜力。

原文摘要 · Abstract (English)

We introduce native-resolution image synthesis, a novel generative modeling paradigm that enables the synthesis of images at arbitrary resolutions and aspect ratios. This approach overcomes the limitations of conventional fixed-resolution, square-image methods by natively handling variable-length visual tokens, a core challenge for traditional techniques. To this end, we introduce the Native-resolution diffusion Transformer (NiT), an architecture designed to explicitly model varying resolutions and aspect ratios within its denoising process. Free from the constraints of fixed formats, NiT learns intrinsic visual distributions from images spanning a broad range of resolutions and aspect ratios. Notably, a single NiT model simultaneously achieves the state-of-the-art performance on both ImageNet-256x256 and 512x512 benchmarks. Surprisingly, akin to the robust zero-shot capabilities seen in advanced large language models, NiT, trained solely on ImageNet, demonstrates excellent zero-shot generalization performance. It successfully generates high-fidelity images at previously unseen high resolutions (e.g., 1536 x 1536) and diverse aspect ratios (e.g., 16:9, 3:1, 4:3), as shown in Figure 1. These findings indicate the significant potential of native-resolution modeling as a bridge between visual generative modeling and advanced LLM methodologies.

图像生成扩散模型多尺度零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。