用高斯混合采样提升长提示图像生成多样性,不损失语义。
PromptMoG: Enhancing Diversity in Long-Prompt Image Generation via Prompt Embedding Mixture-of-Gaussian Sampling
- 在提示嵌入空间用高斯混合采样,增强生成多样性。
- 在四个主流模型上验证,多样性显著提升且语义不变。
- 提出LPD-Bench基准,可统一评估长提示生成的保真与多样。
近期基于大规模修正流模型的文本到图像生成取得了显著视觉效果,但长提示下的模型行为仍缺乏深入研究。长提示虽能增强图像保真度,却常抑制多样性,导致输出重复、创意不足。本文系统研究该保真-多样性权衡问题,发现主流模型随提示长度增加,多样性明显下降。为此,我们构建了用于评估长提示生成保真与多样性的基准LPD-Bench。基于分析,提出理论框架通过提示重述提升采样熵,并设计无需训练的PromptMoG方法:在嵌入空间中从高斯混合模型采样提示嵌入,以增强多样性并保持语义。在SD3.5-Large、Flux.1-Krea-Dev、CogView4和Qwen-Image四个先进模型上的大量实验表明,PromptMoG能持续提升长提示生成的多样性,且无语义漂移。
原文摘要 · Abstract (English)
Recent advances in text-to-image (T2I) generation have achieved remarkable visual outcomes through large-scale rectified flow models. However, how these models behave under long prompts remains underexplored. Long prompts encode rich content, spatial, and stylistic information that enhances fidelity but often suppresses diversity, leading to repetitive and less creative outputs. In this work, we systematically study this fidelity-diversity dilemma and reveal that state-of-the-art models exhibit a clear drop in diversity as prompt length increases. To enable consistent evaluation, we introduce LPD-Bench, a benchmark designed for assessing both fidelity and diversity in long-prompt generation. Building on our analysis, we develop a theoretical framework that increases sampling entropy through prompt reformulation and propose a training-free method, PromptMoG, which samples prompt embeddings from a Mixture-of-Gaussians in the embedding space to enhance diversity while preserving semantics. Extensive experiments on four state-of-the-art models, SD3.5-Large, Flux.1-Krea-Dev, CogView4, and Qwen-Image, demonstrate that PromptMoG consistently improves long-prompt generation diversity without semantic drifting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。