构建2.6万张服装草图数据集,助力AI精准生成时尚图像
GarmentSketch: Large-scale Sketch-to-Fashion Benchmark

- 用多模态大模型+人工校正生成高质量草图-文本配对数据
- 涵盖21类服装,共26,249对草图与详细描述
- 为设计师与AI协作提供可量化评估的基准数据集
服装草图是设计流程的核心,能快速呈现创意概念。然而,草图驱动的时尚图像生成进展受限于缺乏大规模、高质量的成对数据。为此,我们提出GarmentSketch,一个包含26,249张跨21类服装的时尚草图数据集,每张草图均配有详尽的文本描述。描述通过多阶段流程生成,融合多个多模态大语言模型(MLLMs)与人工介入优化,确保语义准确性和描述丰富性。我们在先进生成模型上对GarmentSketch进行基准测试,提供草图引导的文本到图像生成基线性能。实验揭示了现有方法的潜力与当前局限。通过提供全面且丰富标注的资源,GarmentSketch为提升草图理解、细粒度时尚图像生成及设计中的人机协同奠定了基础。数据集将公开获取:https://khangbdd.github.io/garmentsketch。
原文摘要 · Abstract (English)
Fashion sketching is a cornerstone of design workflows, allowing rapid visualization of creative concepts prior to physical prototyping. Yet, progress in sketch-based fashion image synthesis has been hindered by the absence of large-scale, high-quality paired resources. To bridge this gap, we present GarmentSketch, a novel dataset comprising 26,249 fashion sketches across 21 garment categories, each paired with detailed textual descriptions. Captions were produced through a multi-stage pipeline that integrates multiple multimodal large language models (MLLMs) with human-in-the-loop refinement, ensuring both semantic accuracy and descriptive richness. We benchmark GarmentSketch on state-of-the-art generative models, providing baseline performance for sketch-guided text-to-image generation. Our experiments reveal both the promise and the current limitations of existing methods. By offering a comprehensive and richly annotated resource, GarmentSketch establishes a foundation for advancing sketch understanding, fine-grained fashion image generation, and creative human-AI collaboration in design. The dataset will be available at: https://khangbdd.github.io/garmentsketch.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。