arXiv:2505.10046cs.CV2025-05CVPR被引 16

探索大模型与扩散Transformer深度融合的文本生成图像方法

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis

  • 系统对比不同融合方式,分析关键设计选择
  • 提供可复现的大规模训练方案,提升生成效果
  • 适合关注多模态生成与模型融合的研究者

本文不提出新方法,而是深入探索文本到图像生成中一个重要但未被充分研究的设计空间——即大语言模型(LLMs)与扩散Transformer(DiTs)的深度融合。以往研究多关注整体性能,缺乏对替代方法的细致比较,且关键设计细节与训练策略常被省略,导致该方法真实潜力存在不确定性。为填补这些空白,我们开展了一项实证研究,通过与基准方法进行受控对比,分析关键设计选择,并提供清晰、可复现的规模化训练方案。本工作旨在为未来多模态生成研究提供可靠数据支持与实用指导。

原文摘要 · Abstract (English)

This paper does not describe a new method; instead, it provides a thorough exploration of an important yet understudied design space related to recent advances in text-to-image synthesis -- specifically, the deep fusion of large language models (LLMs) and diffusion transformers (DiTs) for multi-modal generation. Previous studies mainly focused on overall system performance rather than detailed comparisons with alternative methods, and key design details and training recipes were often left undisclosed. These gaps create uncertainty about the real potential of this approach. To fill these gaps, we conduct an empirical study on text-to-image generation, performing controlled comparisons with established baselines, analyzing important design choices, and providing a clear, reproducible recipe for training at scale. We hope this work offers meaningful data points and practical guidelines for future research in multi-modal generation.

文本生成图像多模态生成模型融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。