arXiv:2504.14011cs.CVcs.AI2025-04被引 5

用文字描述衣服,AI自动找相似款并生成定制化穿搭图

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation

  • 通过检索匹配文本描述的服装,融合多件参考图生成新图像
  • 在Dress Code数据集上优于现有方法,细节还原更精准
  • 适合想快速试穿不同风格但无实物或草图的用户

近年来,时尚产业日益采用AI技术提升用户体验,推动电商与虚拟应用发展。其中,虚拟试穿和多模态时尚图像编辑(利用文本、服装草图、身体姿态等输入)成为研究热点。扩散模型因其出色的图像质量和多样性,已成为主流生成方法。然而,多数现有虚拟试穿方法依赖特定服装输入,在实际场景中难以满足用户仅提供文字描述的需求。为此,本文提出时尚检索增强生成(Fashion-RAG),一种基于文本描述定制时尚单品的新方法。该方法先检索与输入描述匹配的多件服装,再通过属性融合生成个性化图像。我们采用文本反转技术,将检索到的服装图像映射至Stable Diffusion文本编码器的嵌入空间,实现检索内容与生成过程的无缝结合。在Dress Code数据集上的实验表明,Fashion-RAG在定性和定量评估中均优于现有方法,能有效捕捉检索服装的细粒度视觉特征。据我们所知,这是首个专为多模态时尚图像编辑设计的检索增强生成框架。

原文摘要 · Abstract (English)

In recent years, the fashion industry has increasingly adopted AI technologies to enhance customer experience, driven by the proliferation of e-commerce platforms and virtual applications. Among the various tasks, virtual try-on and multimodal fashion image editing -- which utilizes diverse input modalities such as text, garment sketches, and body poses -- have become a key area of research. Diffusion models have emerged as a leading approach for such generative tasks, offering superior image quality and diversity. However, most existing virtual try-on methods rely on having a specific garment input, which is often impractical in real-world scenarios where users may only provide textual specifications. To address this limitation, in this work we introduce Fashion Retrieval-Augmented Generation (Fashion-RAG), a novel method that enables the customization of fashion items based on user preferences provided in textual form. Our approach retrieves multiple garments that match the input specifications and generates a personalized image by incorporating attributes from the retrieved items. To achieve this, we employ textual inversion techniques, where retrieved garment images are projected into the textual embedding space of the Stable Diffusion text encoder, allowing seamless integration of retrieved elements into the generative process. Experimental results on the Dress Code dataset demonstrate that Fashion-RAG outperforms existing methods both qualitatively and quantitatively, effectively capturing fine-grained visual details from retrieved garments. To the best of our knowledge, this is the first work to introduce a retrieval-augmented generation approach specifically tailored for multimodal fashion image editing.

时尚生成多模态检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。