构建13万规模电商时尚数据集,提升穿搭生成的视觉一致性。
Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition

- 设计统一多模态条件框架,融合文本与图像提示
- 在真实场景中生成效果优于现有最先进方法
- 适合做多模态生成、时尚内容创作的研究者
当前时尚穿搭生成研究多关注通过参考图像和文本提示中的关键信息提升服装视觉一致性,但该领域潜力尚未充分挖掘,亟需全面的电商数据集及对多模态条件的精细利用。本文提出一个全新的电商时尚数据集Fashion130k,涵盖多种场合、模特和服装类型,共13万条样本。为实现服装的一致生成,设计了统一多模态条件(UMC)框架,将文本与视觉提示对齐并融合。具体地,引入嵌入精炼模块,其中提出融合变压器(Fusion Transformer)以调整文本与图像之间的模态差异,生成统一嵌入。基于此,重设计生成模型中的注意力机制,强化提示与噪声图像间的关联,使噪声图像能选择提示中的关键词元,从而实现一致穿搭生成。大量真实应用与基准测试表明,该框架在视觉一致性上表现优异,优于现有最先进方法。
原文摘要 · Abstract (English)
Recent research work on fashion outfit generation focuses on promoting visual consistency of garments by leveraging key information from reference image and text prompt. However, the potential of outfit generation remains underexplored, requiring comprehensive e-commercial dataset and elaborative utilization of multi-modal condition. In this paper, we propose a brand-new e-commerce dataset, named Fashion130k, with various occasions, models, and garment types. For the consistent generation of garment, we design a framework with Unified Multi-modal Condition (UMC) to align and integrate the text and visual prompts into generation model. Specifically, we explore an embedding refiner to extract the unified embeddings of multi-modal prompts, within which a Fusion Transformer is proposed to align the multi-modal embeddings by adjusting the modality gap between text and image. Based on unified embeddings, the attention in generation model is redesigned to emphasis the correlations between prompts and noise image, inducing that the noise image can select the pivotal tokens of prompts for consistent outfit generation. Our dataset and proposed framework offer a general and nuanced exploration of multi-modal prompts for generation models. Extensive experiments on real-world applications and benchmark demonstrate the effectiveness of UMC in visual consistency, achieving promising result than that of SoTA methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。