arXiv:2510.00633cs.CVcs.LG2025-10

构建首个成衣时尚大片数据集,推动虚拟拍摄从静态展示走向创意叙事。

Virtual Fashion Photo-Shoots: Building a Large-Scale Garment-Lookbook Dataset

  • 通过视觉-语言推理与目标定位,自动对齐不同领域的成衣与杂志图像。
  • 建成包含10,000对高质量、50,000对中等质量、300,000对低质量的成衣-画册配对数据集。
  • 适合关注时尚生成、创意影像合成与跨域图像对齐的研究者使用。

时尚图像生成目前多聚焦于虚拟试穿等狭窄任务,通常在干净的摄影棚环境中呈现衣物。而时尚编辑则通过动态姿势、多样场景和精心设计的视觉叙事展现服饰。本文提出虚拟时尚大片任务,旨在将标准化成衣图像转化为具有情境意义的时尚编辑影像。为支持这一新方向,我们构建了首个大规模成衣-画册配对数据集,弥合电商与时尚媒体之间的差距。由于此类配对数据难以获取,我们设计了一种自动化检索管道,结合视觉-语言推理与对象级定位实现跨域对齐。最终构建的数据集包含三个质量等级:高质量(10,000对)、中等质量(50,000对)和低质量(300,000对)。该数据集为模型突破目录式生成、迈向体现创造力、氛围感与故事性的时尚图像提供了基础。

原文摘要 · Abstract (English)

Fashion image generation has so far focused on narrow tasks such as virtual try-on, where garments appear in clean studio environments. In contrast, editorial fashion presents garments through dynamic poses, diverse locations, and carefully crafted visual narratives. We introduce the task of virtual fashion photo-shoot, which seeks to capture this richness by transforming standardized garment images into contextually grounded editorial imagery. To enable this new direction, we construct the first large-scale dataset of garment-lookbook pairs, bridging the gap between e-commerce and fashion media. Because such pairs are not readily available, we design an automated retrieval pipeline that aligns garments across domains, combining visual-language reasoning with object-level localization. We construct a dataset with three garment-lookbook pair accuracy levels: high quality (10,000 pairs), medium quality (50,000 pairs), and low quality (300,000 pairs). This dataset offers a foundation for models that move beyond catalog-style generation and toward fashion imagery that reflects creativity, atmosphere, and storytelling.

时尚生成图像对齐数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。