arXiv:2505.21062cs.CV2025-05中稿 · ICLR被引 7

从穿衣服的人像生成多品类服装产品图,解决细节丢失问题。

Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals

  • 用图文掩码联合建模,消除单图视觉歧义
  • 多类别服装生成质量显著提升,细节更清晰
  • 适合电商、数据集构建与大模型训练

虚拟试穿(VTON)已广泛用于将衣物渲染到人像上,而其逆任务——虚拟脱衣(VTOFF)却很少被关注。VTOFF旨在从穿着衣物的人像照片中恢复标准化的服装产品图像,对电商平台、大规模数据集构建及基础模型训练具有重要应用价值。与需处理多样姿态和风格的VTON不同,VTOFF天然受益于输出格式的一致性,即平铺的服装图像。然而现有方法存在两大局限:(i) 仅依赖单张照片的视觉线索常导致歧义;(ii) 生成图像往往损失精细细节,影响实际应用。为此,我们提出TEMU-VTOFF,一种文本增强的多类别框架。该架构基于双DiT骨干网络,结合多模态注意力机制,联合利用图像、文本和掩码信息以消除视觉歧义并实现跨类别鲁棒特征学习。为明确缓解细节退化,我们进一步设计了对齐模块,精修服装结构与纹理,确保高质量输出。在VITON-HD和Dress Code上的大量实验表明,TEMU-VTOFF达到新最优性能,显著提升视觉真实感与与目标服装的一致性。

原文摘要 · Abstract (English)

Virtual try-on (VTON) has been widely explored for rendering garments onto person images, while its inverse task, virtual try-off (VTOFF), remains largely overlooked. VTOFF aims to recover standardized product images of garments directly from photos of clothed individuals. This capability is of great practical importance for e-commerce platforms, large-scale dataset curation, and the training of foundation models. Unlike VTON, which must handle diverse poses and styles, VTOFF naturally benefits from a consistent output format in the form of flat garment images. However, existing methods face two major limitations: (i) exclusive reliance on visual cues from a single photo often leads to ambiguity, and (ii) generated images usually suffer from loss of fine details, limiting their real-world applicability. To address these challenges, we introduce TEMU-VTOFF, a Text-Enhanced MUlti-category framework for VTOFF. Our architecture is built on a dual DiT-based backbone equipped with a multimodal attention mechanism that jointly exploits image, text, and mask information to resolve visual ambiguities and enable robust feature learning across garment categories. To explicitly mitigate detail degradation, we further design an alignment module that refines garment structures and textures, ensuring high-quality outputs. Extensive experiments on VITON-HD and Dress Code show that TEMU-VTOFF achieves new state-of-the-art performance, substantially improving both visual realism and consistency with target garments.

虚拟试穿图像生成多模态服装生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。