通过分层对齐与扩散机制,提升文本生成图像的语义准确性和视觉质量。
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models
- 分两阶段处理文本:全局与局部特征分离,增强语义对齐
- 多阶段扩散过程生成高保真图像,在双评估集上超越现有方法
- 适合需要精确图文一致性的复杂场景生成任务
文本到图像生成在引入大视觉语言模型后取得显著进展,但仍面临复杂文本描述与高质量、视觉连贯图像对齐的挑战。本文提出视觉-语言对齐扩散模型(VLAD),采用双流策略结合语义对齐与分层扩散。VLAD使用上下文组合模块(CCM)将文本提示分解为全局与局部表示,确保与视觉特征精准匹配。同时,引入多阶段扩散过程与分层引导机制,生成高保真图像。在MARIO-Eval和INNOVATOR-Eval基准测试中,VLAD在图像质量、语义对齐和文本渲染准确性方面均显著优于当前最优方法。人工评估进一步验证其优越性能,使其成为复杂场景下文本到图像生成的有力方案。
原文摘要 · Abstract (English)
Text-to-image generation has witnessed significant advancements with the integration of Large Vision-Language Models (LVLMs), yet challenges remain in aligning complex textual descriptions with high-quality, visually coherent images. This paper introduces the Vision-Language Aligned Diffusion (VLAD) model, a generative framework that addresses these challenges through a dual-stream strategy combining semantic alignment and hierarchical diffusion. VLAD utilizes a Contextual Composition Module (CCM) to decompose textual prompts into global and local representations, ensuring precise alignment with visual features. Furthermore, it incorporates a multi-stage diffusion process with hierarchical guidance to generate high-fidelity images. Experiments conducted on MARIO-Eval and INNOVATOR-Eval benchmarks demonstrate that VLAD significantly outperforms state-of-the-art methods in terms of image quality, semantic alignment, and text rendering accuracy. Human evaluations further validate the superior performance of VLAD, making it a promising approach for text-to-image generation in complex scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。