arXiv:2502.14373cs.CV2025-02IJCAI被引 3

通过三区先验模拟人类逻辑推理,提升跨品类虚拟试穿效果。

CrossVTON: Mimicking the Logic Reasoning on Cross-category Virtual Try-on guided by Tri-zone Priors

  • 将人体图像分为试穿、重建和想象三区,分区域处理服装适配。
  • 在跨品类试穿任务上超越现有方法,生成更真实自然的图像。
  • 适合需要高精度跨品类虚拟试穿的电商与设计场景。

尽管基于图像的虚拟试穿系统取得了显著进展,但在跨品类场景下生成逼真且鲁棒的试穿图像仍是挑战。主要难点在于缺乏类人推理能力,即无法有效解决服装与模特尺寸不匹配问题,也难以识别并利用模型图像中各区域的差异化功能。为此,我们借鉴人类认知过程,将复杂的跨品类试穿推理分解为结构化框架,将模型图像划分为试穿区、重建区和想象区,每个区域承担特定角色以适配服装并实现真实合成。为增强模型在跨品类场景下的推理能力,提出一种迭代式数据构造器,涵盖同类别试穿、任意服装转连衣裙、连衣裙转任意服装等多种场景。利用生成的数据集,引入三区先验生成器,通过分析输入服装与模特图像的对齐预期,智能预测三区位置。在三区先验引导下,所提方法CrossVTON在定性和定量评估中均达到领先性能,尤其在跨品类试穿任务上表现优异,满足真实应用场景的复杂需求。

原文摘要 · Abstract (English)

Despite remarkable progress in image-based virtual try-on systems, generating realistic and robust fitting images for cross-category virtual try-on remains a challenging task. The primary difficulty arises from the absence of human-like reasoning, which involves addressing size mismatches between garments and models while recognizing and leveraging the distinct functionalities of various regions within the model images. To address this issue, we draw inspiration from human cognitive processes and disentangle the complex reasoning required for cross-category try-on into a structured framework. This framework systematically decomposes the model image into three distinct regions: try-on, reconstruction, and imagination zones. Each zone plays a specific role in accommodating the garment and facilitating realistic synthesis. To endow the model with robust reasoning capabilities for cross-category scenarios, we propose an iterative data constructor. This constructor encompasses diverse scenarios, including intra-category try-on, any-to-dress transformations (replacing any garment category with a dress), and dress-to-any transformations (replacing a dress with another garment category). Utilizing the generated dataset, we introduce a tri-zone priors generator that intelligently predicts the try-on, reconstruction, and imagination zones by analyzing how the input garment is expected to align with the model image. Guided by these tri-zone priors, our proposed method, CrossVTON, achieves state-of-the-art performance, surpassing existing baselines in both qualitative and quantitative evaluations. Notably, it demonstrates superior capability in handling cross-category virtual try-on, meeting the complex demands of real-world applications.

虚拟试穿图像生成跨品类三区先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。