arXiv:2411.13901cs.CV2024-11中稿 · as a Conference Pa…被引 3

首个时尚文本-穿搭数据集,助力AI精准生成风格化服装设计。

Dressing the Imagination: A Dataset for AI-Powered Translation of Text into Fashion Outfits and A Novel NeRA Adapter for Enhanced Feature Adaptation

  • 构建4330对图文匹配的时尚数据集FLORA,含专业设计术语。
  • 在FLORA上微调模型,生成图像的风格准确率显著提升。
  • 提出NeRA适配器,比传统方法更快更准地捕捉语义关系。

为推动AI驱动的时尚设计发展,我们提出了FLORA(Fashion Language Outfit Representation for Apparel Generation),首个包含4,330组精心标注的时尚穿搭与对应文本描述的综合性数据集。每条描述采用专业设计师常用的行业术语和行话,精准刻画服饰细节与风格要素,有助于生成高保真度的时尚设计。实验表明,在FLORA上微调生成模型可显著提升其根据文字描述生成准确且风格丰富的图像的能力。此外,我们提出一种新型适配器结构NeRA(Nonlinear low-rank Expressive Representation Adapter),基于Kolmogorov-Arnold Networks(KAN),使用可学习的样条非线性变换,替代传统MLP适配器,有效建模复杂语义关系,实现更强的保真度、更快收敛速度和更好的语义对齐。在FLORA和LAION-5B数据集上的大量实验验证了NeRA的优势。我们将开源FLORA数据集及代码实现。

原文摘要 · Abstract (English)

Specialized datasets that capture the fashion industry's rich language and styling elements can boost progress in AI-driven fashion design. We present FLORA, (Fashion Language Outfit Representation for Apparel Generation), the first comprehensive dataset containing 4,330 curated pairs of fashion outfits and corresponding textual descriptions. Each description utilizes industry-specific terminology and jargon commonly used by professional fashion designers, providing precise and detailed insights into the outfits. Hence, the dataset captures the delicate features and subtle stylistic elements necessary to create high-fidelity fashion designs. We demonstrate that fine-tuning generative models on the FLORA dataset significantly enhances their capability to generate accurate and stylistically rich images from textual descriptions of fashion sketches. FLORA will catalyze the creation of advanced AI models capable of comprehending and producing subtle, stylistically rich fashion designs. It will also help fashion designers and end-users to bring their ideas to life. As a second orthogonal contribution, we introduce NeRA (Nonlinear low-rank Expressive Representation Adapter), a novel adapter architecture based on Kolmogorov-Arnold Networks (KAN). Unlike traditional PEFT techniques such as LoRA, LoKR, DoRA, and LoHA that use MLP adapters, NeRA uses learnable spline-based nonlinear transformations, enabling superior modeling of complex semantic relationships, achieving strong fidelity, faster convergence and semantic alignment. Extensive experiments on our proposed FLORA and LAION-5B datasets validate the superiority of NeRA over existing adapters. We will open-source both the FLORA dataset and our implementation code.

时尚生成数据集适配器AI设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。