arXiv:2411.14808cs.CVcs.AI2024-11被引 4

用连续令牌预测生成4K级逼真图像,突破自回归模型瓶颈

High-Resolution Image Synthesis via Next-Token Prediction

  • 基于去噪联合嵌入架构,融合文本与视觉特征生成图像
  • 引入流匹配损失和旋转位置编码,实现任意分辨率连续学习
  • 动态数据反馈机制提升模型对难样本的生成能力

近期自回归模型在类别条件图像生成中表现优异,但针对高分辨率文本到图像生成的下一令牌预测研究仍不充分。本文提出基于连续令牌的自回归模型 D-JEPA·T2I,通过架构与训练策略创新,实现最高4K分辨率的高质量、逼真图像合成。架构上采用去噪联合嵌入预测架构(D-JEPA),结合多模态视觉变换器有效融合文本与视觉特征;引入流匹配损失与视觉旋转位置编码(VoPE),支持连续分辨率学习。训练策略上提出数据反馈机制,基于统计分析动态调整采样过程,并引入在线学习判别器模型,促使模型跳出舒适区,减少对已掌握场景的冗余训练,聚焦解决生成质量较差的挑战性案例。首次实现基于下一令牌预测的顶尖高分辨率图像合成。

原文摘要 · Abstract (English)

Recently, autoregressive models have demonstrated remarkable performance in class-conditional image generation. However, the application of next-token prediction to high-resolution text-to-image generation remains largely unexplored. In this paper, we introduce \textbf{D-JEPA$\cdot$T2I}, an autoregressive model based on continuous tokens that incorporates innovations in both architecture and training strategy to generate high-quality, photorealistic images at arbitrary resolutions, up to 4K. Architecturally, we adopt the denoising joint embedding predictive architecture (D-JEPA) while leveraging a multimodal visual transformer to effectively integrate textual and visual features. Additionally, we introduce flow matching loss alongside the proposed Visual Rotary Positional Embedding (VoPE) to enable continuous resolution learning. In terms of training strategy, we propose a data feedback mechanism that dynamically adjusts the sampling procedure based on statistical analysis and an online learning critic model. This encourages the model to move beyond its comfort zone, reducing redundant training on well-mastered scenarios and compelling it to address more challenging cases with suboptimal generation quality. For the first time, we achieve state-of-the-art high-resolution image synthesis via next-token prediction.

图像生成自回归模型4K合成连续令牌

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。