arXiv:2501.15891cs.CV2025-01ICCV被引 68

无需遮罩或姿态,一句话指令即可实现任意服装试穿生成。

Any2AnyTryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing Tasks

  • 用自适应位置嵌入提升模型对不同尺寸服装的适配能力。
  • 构建了最大开源服装试穿数据集LAION-Garment,含超10万对图像。
  • 支持文本指令控制,可灵活生成无遮罩、无姿态约束的试穿图。

基于图像的虚拟试穿(VTON)旨在将输入服装图像迁移至目标人物图像上生成试穿结果。然而,成对服装-模特数据稀缺导致现有方法在泛化性和质量上受限,且难以生成无遮罩试穿。为解决数据稀缺问题,如Stable Garment和MMTryon等方法采用合成数据策略,有效增加模型侧成对数据量。但现有方法通常仅限于特定试穿任务,缺乏用户友好性。为此,我们提出Any2AnyTryon,可根据不同文本指令和模特服装图像生成试穿结果,摆脱对遮罩、姿态等条件的依赖。具体而言,我们构建了目前最大的开源服装试穿数据集LAION-Garment(含超10万对图像)。引入自适应位置嵌入,使模型能根据不同尺寸与类别的输入图像生成高质量的试穿图像,显著提升泛化与可控性。实验表明,Any2AnyTryon在灵活性、可控性与生成质量上均优于现有方法。

原文摘要 · Abstract (English)

Image-based virtual try-on (VTON) aims to generate a virtual try-on result by transferring an input garment onto a target person's image. However, the scarcity of paired garment-model data makes it challenging for existing methods to achieve high generalization and quality in VTON. Also, it limits the ability to generate mask-free try-ons. To tackle the data scarcity problem, approaches such as Stable Garment and MMTryon use a synthetic data strategy, effectively increasing the amount of paired data on the model side. However, existing methods are typically limited to performing specific try-on tasks and lack user-friendliness. To enhance the generalization and controllability of VTON generation, we propose Any2AnyTryon, which can generate try-on results based on different textual instructions and model garment images to meet various needs, eliminating the reliance on masks, poses, or other conditions. Specifically, we first construct the virtual try-on dataset LAION-Garment, the largest known open-source garment try-on dataset. Then, we introduce adaptive position embedding, which enables the model to generate satisfactory outfitted model images or garment images based on input images of different sizes and categories, significantly enhancing the generalization and controllability of VTON generation. In our experiments, we demonstrate the effectiveness of our Any2AnyTryon and compare it with existing methods. The results show that Any2AnyTryon enables flexible, controllable, and high-quality image-based virtual try-on generation. https://logn-2024.github.io/Any2anyTryonProjectPage

虚拟试穿图像生成自适应嵌入文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。