构建多视角时尚数据集,支持虚拟试穿与尺码估计。
MV-Fashion: Towards Enabling Virtual Try-On and Size Estimation with Multi-View Paired Data
- 基于80人3273段视频,捕捉真实衣物动态与多层穿搭。
- 提供像素级标注、材质属性及配对的穿着/平铺图像。
- 适合研究虚拟试穿、尺码估算和新视角生成的团队使用。
现有4D人体数据集在时尚研究中存在不足,缺乏逼真的服装动态或任务特定标注。合成数据集存在真实性差距,而真实采集数据又缺少详细标注和配对数据,难以支撑虚拟试穿(VTON)与尺码估计任务。为此,我们提出MV-Fashion,一个面向时尚分析的大规模多视角视频数据集。该数据集包含80名不同被试、每人3-10套服饰,共3,273个序列,总计7250万帧。其设计聚焦于捕捉复杂真实的服装动态,如多层叠加、卷袖、扎衣等多样化穿搭。核心贡献在于丰富的数据表示:包含像素级语义标注、真实的材料属性(如弹性)、以及3D点云。关键优势在于提供配对数据——同步的多视角穿着视频与对应的平铺商品图。我们利用该数据集建立时尚任务基准,涵盖虚拟试穿、服装尺码估计与新视角合成。数据集已公开:https://hunorlaczko.github.io/MV-Fashion。
原文摘要 · Abstract (English)
Existing 4D human datasets fall short for fashion-specific research, lacking either realistic garment dynamics or task-specific annotations. Synthetic datasets suffer from a realism gap, whereas real-world captures lack the detailed annotations and paired data required for virtual try-on (VTON) and size estimation tasks. To bridge this gap, we introduce MV-Fashion, a large-scale, multi-view video dataset engineered for domain-specific fashion analysis. MV-Fashion features 3,273 sequences (72.5 million frames) from 80 diverse subjects wearing 3-10 outfits each. It is designed to capture complex, real-world garment dynamics, including multiple layers and varied styling (e.g. rolled sleeves, tucked shirt). A core contribution is a rich data representation that includes pixel-level semantic annotations, ground-truth material properties like elasticity, and 3D point clouds. Crucially for VTON applications, MV-Fashion provides paired data: multi-view synchronized captures of worn garments alongside their corresponding flat, catalogue images. We leverage this dataset to establish baselines for fashion-centric tasks, including virtual try-on, clothing size estimation, and novel view synthesis. The dataset is available at https://hunorlaczko.github.io/MV-Fashion .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。