用大模型动态调整扩散模型,让图文生成更准更快。
Weak Supervision Dynamic KL-Weighted Diffusion Models Guided by Large Language Models
- 结合大模型语义理解与动态KL加权优化生成过程
- 在COCO数据集上显著提升图像真实感与文本对齐度
- 适合需要高质量图文生成的科研与工业场景
本文提出一种融合大语言模型(LLMs)与扩散模型的新方法,以提升文本到图像生成的质量与效率。通过引入动态KL权重策略优化扩散过程,并利用预训练大模型增强语义理解,引导图像生成。该方法有效缓解了计算低效、训练不稳定及对文本变化鲁棒性差等问题。在COCO数据集上的实验表明,其性能优于传统GAN模型,无论在定量指标还是定性评估中均表现更优。消融实验与人工评价进一步验证了其在图像真实性、文本相关性和整体美学质量方面的优势。该方法在其他多模态任务中也展现出良好的可扩展性,具备广泛的应用潜力。
原文摘要 · Abstract (English)
In this paper, we presents a novel method for improving text-to-image generation by combining Large Language Models (LLMs) with diffusion models, a hybrid approach aimed at achieving both higher quality and efficiency in image synthesis from text descriptions. Our approach introduces a new dynamic KL-weighting strategy to optimize the diffusion process, along with incorporating semantic understanding from pre-trained LLMs to guide the generation process. The proposed method significantly improves both the visual quality and alignment of generated images with text descriptions, addressing challenges such as computational inefficiency, instability in training, and robustness to textual variability. We evaluate our method on the COCO dataset and demonstrate its superior performance over traditional GAN-based models, both quantitatively and qualitatively. Extensive experiments, including ablation studies and human evaluations, confirm that our method outperforms existing approaches in terms of image realism, relevance to the input text, and overall aesthetic quality. Our approach also shows promise in scalability to other multimodal tasks, making it a versatile solution for a wide range of generative applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。