用大模型和扩散模型生成跨文化时尚设计,解决偏见与语图不符问题。
Cross-Cultural Fashion Design via Interactive Large Language Models and Diffusion Models
- 用大模型优化文本提示,提升语义准确性
- 在增强版DeepFashion+数据集上实现更低FID、更高IS
- 适合想做跨文化时尚生成的研究者与设计师
时尚内容生成是人工智能与创意设计交叉的新兴领域,应用涵盖虚拟试穿和跨文化设计原型。现有方法常面临文化偏见、可扩展性差及文本提示与生成图像对齐不足的问题,尤其在弱监督条件下表现不佳。本文提出一种新框架,将大语言模型(LLMs)与潜在扩散模型(LDMs)结合。该方法利用LLM对文本提示进行语义优化,并引入弱监督过滤模块,有效利用噪声或弱标注数据。通过在增强版DeepFashion+数据集上微调LDM,所提方法达到当前最优性能。实验表明,其显著优于基线,实现更低的弗雷切特起始距离(FID)和更高的起始分数(IS),且人工评估确认其能生成具有文化多样性和语义相关性的时尚内容。结果表明,大模型引导的扩散模型在推动可扩展、包容性AI时尚创新方面具有潜力。
原文摘要 · Abstract (English)
Fashion content generation is an emerging area at the intersection of artificial intelligence and creative design, with applications ranging from virtual try-on to culturally diverse design prototyping. Existing methods often struggle with cultural bias, limited scalability, and alignment between textual prompts and generated visuals, particularly under weak supervision. In this work, we propose a novel framework that integrates Large Language Models (LLMs) with Latent Diffusion Models (LDMs) to address these challenges. Our method leverages LLMs for semantic refinement of textual prompts and introduces a weak supervision filtering module to effectively utilize noisy or weakly labeled data. By fine-tuning the LDM on an enhanced DeepFashion+ dataset enriched with global fashion styles, the proposed approach achieves state-of-the-art performance. Experimental results demonstrate that our method significantly outperforms baselines, achieving lower Frechet Inception Distance (FID) and higher Inception Scores (IS), while human evaluations confirm its ability to generate culturally diverse and semantically relevant fashion content. These results highlight the potential of LLM-guided diffusion models in driving scalable and inclusive AI-driven fashion innovation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。