构建百万级服装合身度数据集,让虚拟试穿更真实反映尺寸差异。
FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-On

- 用3D物理模拟生成逼真合身效果的合成服装图像
- 包含113万组配对图像,精确标注人体与服装尺寸
- 适合关注服装合身度、虚拟试穿真实性的研究者
虚拟试穿(VTO)旨在将衣物图像合成到人物图像上,保持其原始姿态和身份。尽管现有方法在呈现衣物外观方面表现优异,但普遍忽视了试穿体验中的关键因素——合身度,例如大号衬衫穿在小个子身上时的真实效果。主要障碍在于缺乏提供精确人体与衣物尺寸信息的数据集,尤其缺少显著不合身的案例。因此当前方法默认生成合身结果,忽略实际尺寸差异。本文首次提出解决该问题的方案:构建大规模合身感知虚拟试穿数据集FIT,包含超过113万组试穿图像三元组及精准的体形与衣物测量数据。通过可扩展的合成策略实现:(1)使用GarmentCode生成3D衣物并经物理模拟披挂,捕捉真实合身效果;(2)引入新颖的再贴图框架,将合成渲染转化为写实图像,严格保留几何结构;(3)在再贴图模型中融入身份保持机制,生成同一人穿不同衣物的成对图像,支持监督训练。最后,基于FIT数据集训练出基准合身感知虚拟试穿模型。我们的数据与成果为合身感知虚拟试穿树立了新标准,并为未来研究提供可靠基准。所有数据与代码将在项目页面公开:https://johannakarras.github.io/FIT。
原文摘要 · Abstract (English)
Given a person and a garment image, virtual try-on (VTO) aims to synthesize a realistic image of the person wearing the garment, while preserving their original pose and identity. Although recent VTO methods excel at visualizing garment appearance, they largely overlook a crucial aspect of the try-on experience: the accuracy of garment fit -- for example, depicting how an extra-large shirt looks on an extra-small person. A key obstacle is the absence of datasets that provide precise garment and body size information, particularly for "ill-fit" cases, where garments are significantly too large or too small. Consequently, current VTO methods default to generating well-fitted results regardless of the garment or person size. In this paper, we take the first steps towards solving this open problem. We introduce FIT (Fit-Inclusive Try-on), a large-scale VTO dataset comprising over 1.13M try-on image triplets accompanied by precise body and garment measurements. We overcome the challenges of data collection via a scalable synthetic strategy: (1) We programmatically generate 3D garments using GarmentCode and drape them via physics simulation to capture realistic garment fit. (2) We employ a novel re-texturing framework to transform synthetic renderings into photorealistic images while strictly preserving geometry. (3) We introduce person identity preservation into our re-texturing model to generate paired person images (same person, different garments) for supervised training. Finally, we leverage our FIT dataset to train a baseline fit-aware virtual try-on model. Our data and results set the new state-of-the-art for fit-aware virtual try-on, as well as offer a robust benchmark for future research. We will make all data and code publicly available on our project page: https://johannakarras.github.io/FIT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。