arXiv:2603.14153cs.CV2026-03被引 1

首个面向完整穿搭的虚拟试穿数据集,支持多服饰与配饰精细还原。

Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories

  • 构建包含80K穿搭对的多参考数据集,每套穿搭含3-12件服装
  • 现有方法在完整穿搭试穿中易出现错层与图像失真,准确率不足
  • 适合研究穿搭生成、跨模态对齐及虚拟时尚系统的开发者

虚拟试穿(VTON)已能实现单件服饰的可视化,但真实时尚以多件服装、配饰、细分类别、叠穿和多样化搭配为核心,当前系统仍难以处理。现有数据集类别有限,缺乏穿搭多样性。我们提出Garments2Look,首个大规模多模态穿搭级虚拟试穿数据集,包含80,000个「多服饰→一穿搭」样本,覆盖40大类、300多个细分类别。每对样本包括3-12件参考服装(平均4.48件)、穿戴者图像及详尽物品与试穿文本标注。通过启发式构建穿搭列表并结合自动化过滤与人工验证,确保数据真实性与多样性。我们适配前沿VTON与通用图像编辑模型建立基线,结果表明当前方法在完整穿搭试穿中难以实现无缝衔接,常出现层序错误与伪影。

原文摘要 · Abstract (English)

Virtual try-on (VTON) has advanced single-garment visualization, yet real-world fashion centers on full outfits with multiple garments, accessories, fine-grained categories, layering, and diverse styling, remaining beyond current VTON systems. Existing datasets are category-limited and lack outfit diversity. We introduce Garments2Look, the first large-scale multimodal dataset for outfit-level VTON, comprising 80K many-garments-to-one-look pairs across 40 major categories and 300+ fine-grained subcategories. Each pair includes an outfit with 3-12 reference garment images (Average 4.48), a model image wearing the outfit, and detailed item and try-on textual annotations. To balance authenticity and diversity, we propose a synthesis pipeline. It involves heuristically constructing outfit lists before generating try-on results, with the entire process subjected to strict automated filtering and human validation to ensure data quality. To probe task difficulty, we adapt SOTA VTON methods and general-purpose image editing models to establish baselines. Results show current methods struggle to try on complete outfits seamlessly and to infer correct layering and styling, leading to misalignment and artifacts.

虚拟试穿多模态数据穿搭生成图像合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。