通过对比对齐与结构引导,提升图文生成的语义准确性和结构一致性。
High Fidelity Text to Image Generation with Contrastive Alignment and Structural Guidance
- 引入对比学习模块增强文本与图像的跨模态对齐。
- 结合语义布局或边缘草图指导生成,提升图像布局完整度与细节质量。
- 无需增加计算量,适合需要高保真生成的场景。
本文针对现有文本驱动图像生成方法在语义对齐精度与结构一致性方面的性能瓶颈,提出一种高保真图像生成方法,融合文本-图像对比约束与结构引导机制。该方法设计对比学习模块,构建强跨模态对齐约束,提升文本与图像间的语义匹配能力;同时利用语义布局图或边缘草图等结构先验,在空间层级上指导生成器进行结构建模,增强生成图像的布局完整性与细节保真度。整体框架联合优化对比损失、结构一致性损失与语义保留损失,采用多目标监督机制提升生成内容的语义一致性和可控性。在COCO-2014数据集上进行系统实验,并分析嵌入维度、文本长度与结构引导强度的敏感性。定量指标显示,该方法在CLIP Score、FID和SSIM上均优于现有方法。结果表明,该方法有效弥合了语义对齐与结构保真之间的差距,且未增加计算复杂度,具备生成语义清晰、结构完整的高质量图像的能力,为联合文本-图像建模与生成提供了可行技术路径。
原文摘要 · Abstract (English)
This paper addresses the performance bottlenecks of existing text-driven image generation methods in terms of semantic alignment accuracy and structural consistency. A high-fidelity image generation method is proposed by integrating text-image contrastive constraints with structural guidance mechanisms. The approach introduces a contrastive learning module that builds strong cross-modal alignment constraints to improve semantic matching between text and image. At the same time, structural priors such as semantic layout maps or edge sketches are used to guide the generator in spatial-level structural modeling. This enhances the layout completeness and detail fidelity of the generated images. Within the overall framework, the model jointly optimizes contrastive loss, structural consistency loss, and semantic preservation loss. A multi-objective supervision mechanism is adopted to improve the semantic consistency and controllability of the generated content. Systematic experiments are conducted on the COCO-2014 dataset. Sensitivity analyses are performed on embedding dimensions, text length, and structural guidance strength. Quantitative metrics confirm the superior performance of the proposed method in terms of CLIP Score, FID, and SSIM. The results show that the method effectively bridges the gap between semantic alignment and structural fidelity without increasing computational complexity. It demonstrates a strong ability to generate semantically clear and structurally complete images, offering a viable technical path for joint text-image modeling and image generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。