arXiv:2506.00908cs.CV2025-06被引 2

提出双尺度框架,精准对齐服装与人体并保留细节纹理。

DS-VTON: An Enhanced Dual-Scale Coarse-to-Fine Framework for Virtual Try-On

  • 分两阶段生成:先粗后细,先对齐结构再恢复细节。
  • 在多个基准上超越现有方法,结构对齐与纹理保真度均领先。
  • 无需人体分割图,全程免掩码设计,更简洁高效。

尽管已有进展,现有虚拟试穿方法仍难以同时解决两个核心挑战:准确对齐服装图像与目标人体,以及保持服装的细粒度纹理与图案。这两个需求直接对应于粗到细生成范式,其中粗阶段处理结构对齐,细阶段恢复丰富细节。受此启发,我们提出DS-VTON,一种增强型双尺度粗到细框架,更有效地应对试穿问题。该框架包含两个阶段:第一阶段生成低分辨率试穿结果以捕捉服装与人体间的语义对应关系,减少细节有助于实现稳健的结构对齐;第二阶段通过噪声-图像混合的融合精炼扩散过程重构高分辨率输出,通过细化尺度间残差来强调纹理保真度,并有效修正低分辨率阶段的细粒度误差。此外,本方法采用完全无掩码生成策略,摆脱对人体解析图或分割掩码的依赖。大量实验表明,DS-VTON不仅达到当前最优性能,在多个标准虚拟试穿基准上,其结构对齐和纹理保真度始终显著优于先前方法。

原文摘要 · Abstract (English)

Despite recent progress, most existing virtual try-on methods still struggle to simultaneously address two core challenges: accurately aligning the garment image with the target human body, and preserving fine-grained garment textures and patterns. These two requirements map directly onto a coarse-to-fine generation paradigm, where the coarse stage handles structural alignment and the fine stage recovers rich garment details. Motivated by this observation, we propose DS-VTON, an enhanced dual-scale coarse-to-fine framework that tackles the try-on problem more effectively. DS-VTON consists of two stages: the first stage generates a low-resolution try-on result to capture the semantic correspondence between garment and body, where reduced detail facilitates robust structural alignment. In the second stage, a blend-refine diffusion process reconstructs high-resolution outputs by refining the residual between scales through noise-image blending, emphasizing texture fidelity and effectively correcting fine-detail errors from the low-resolution stage. In addition, our method adopts a fully mask-free generation strategy, eliminating reliance on human parsing maps or segmentation masks. Extensive experiments show that DS-VTON not only achieves state-of-the-art performance but consistently and significantly surpasses prior methods in both structural alignment and texture fidelity across multiple standard virtual try-on benchmarks.

虚拟试穿双尺度扩散模型免掩码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。