用12万件服装数据训练多模态模型,一键生成与编辑服装设计。
AIpparel: A Multimodal Foundation Model for Digital Garments
- 基于12万+独特服装的多模态数据,微调大模型生成服装
- 支持文本/图像到服装的生成,还能交互式编辑款式
- 创新编码方案让模型高效学习复杂缝制图
服装对人类至关重要,兼具保护、文化表达与个性展现功能。但设计过程仍高度依赖人工,耗时费力。为此,我们提出AIpparel——一个用于生成与编辑缝制图的多模态基础模型。该模型在自建的超大规模数据集上进行微调,包含超过12万件独特服装,每件均配有文本、图像和缝制图等多模态标注。我们还提出一种新颖的分词方案,能紧凑编码复杂的缝制图,使大语言模型高效学习并预测。AIpparel在单模态任务(如文本到服装、图像到服装)中达到领先水平,并实现交互式服装编辑等全新应用。项目主页:https://georgenakayama.github.io/AIpparel/
原文摘要 · Abstract (English)
Apparel is essential to human life, offering protection, mirroring cultural identities, and showcasing personal style. Yet, the creation of garments remains a time-consuming process, largely due to the manual work involved in designing them. To simplify this process, we introduce AIpparel, a multimodal foundation model for generating and editing sewing patterns. Our model fine-tunes state-of-the-art large multimodal models (LMMs) on a custom-curated large-scale dataset of over 120,000 unique garments, each with multimodal annotations including text, images, and sewing patterns. Additionally, we propose a novel tokenization scheme that concisely encodes these complex sewing patterns so that LLMs can learn to predict them efficiently. AIpparel achieves state-of-the-art performance in single-modal tasks, including text-to-garment and image-to-garment prediction, and enables novel multimodal garment generation applications such as interactive garment editing. The project website is at https://georgenakayama.github.io/AIpparel/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。