首个支持指令编辑的虚拟试穿数据集,让服装修改更智能可控。
Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off
- 构建统一框架下的大规模指令编辑数据集
- 含146,000+样本,覆盖7类编辑类型与3种服饰类别
- 适合想做可控时尚生成的研究者和开发者
虚拟试穿(VTON)与虚拟脱衣(VTOFF)近年在逼真时装合成与衣物重建方面取得显著进展。然而现有数据集仍为静态,缺乏基于指令的可控编辑能力。本文提出首个统一VTON、VTOFF与文本引导服装编辑的大型基准数据集Dress-ED。每个样本包含一件店内服装图像、穿戴该服装的人体图像、其编辑后的对应图像,以及自然语言指令描述期望修改内容。数据集通过融合多模态大模型(MLLM)理解、扩散模型编辑与大语言模型(LLM)验证的全自动流程构建,共包含超过146,000个经验证的四元组,涵盖三种服装类别与七种编辑类型,包括外观(如颜色、图案、材质)与结构(如袖长、领口)修改。基于此,我们进一步提出一种统一的多模态扩散框架,可联合推理语言指令与视觉服装线索,作为指令驱动式VTON与VTOFF的强基线。数据集与代码已公开于:https://github.com/aimagelab/Dress-ED
原文摘要 · Abstract (English)
Recent advances in Virtual Try-On (VTON) and Virtual Try-Off (VTOFF) have greatly improved photo-realistic fashion synthesis and garment reconstruction. However, existing datasets remain static, lacking instruction-driven editing for controllable and interactive fashion generation. In this work, we introduce the Dress Editing Dataset (Dress-ED), the first large-scale benchmark that unifies VTON, VTOFF, and text-guided garment editing within a single framework. Each sample in Dress-ED includes an in-shop garment image, the corresponding person image wearing the garment, their edited counterparts, and a natural-language instruction of the desired modification. Built through a fully automated multimodal pipeline that integrates MLLM-based garment understanding, diffusion-based editing, and LLM-guided verification, Dress-ED comprises over 146k verified quadruplets spanning three garment categories and seven edit types, including both appearance (e.g., color, pattern, material) and structural (e.g., sleeve length, neckline) modifications. Based on this benchmark, we further propose a unified multimodal diffusion framework that jointly reasons over linguistic instructions and visual garment cues, serving as a strong baseline for instruction-driven VTON and VTOFF. Dataset and code available at this link: https://github.com/aimagelab/Dress-ED
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。