arXiv:2607.11233cs.CV2026-07

提出分阶段生成框架,实现快速高保真虚拟试穿。

Structure-Detail Decoupled Autoregressive Generation for Fast and High-Fidelity Virtual Try-On

论文配图:Structure-Detail Decoupled Autoregressive Generation for Fast and High-Fidelity Virtual Try-On
图 1 · 摘自论文原文
  • 先在隐空间合成结构,再在像素空间恢复细节,解耦不同任务。
  • 推理速度比扩散模型快4倍以上,且细节恢复更精准。
  • 可作为通用模块提升现有试穿方法的细节表现。

虚拟试穿(VTON)是需要同时保留人体结构并精确模拟服装形变与细节的双向图像生成问题。基于扩散模型的方法虽能在压缩隐空间中联合建模,但因隐空间压缩导致高频细节丢失,即便采用多步去噪仍难避免。近期视觉自回归(VAR)模型在高速生成方面展现潜力,但因缺乏有效的双向条件机制,尚未应用于VTON。为此,我们提出VAR-VTON,通过服装条件与结构引导实现高效隐空间试穿。然而,隐空间生成仍难以保留细粒度服装细节。我们认为:结构合成(如服装形变、人体布局)适合在隐空间进行,而细粒度细节恢复应放在像素空间处理。基于此,我们进一步提出STAR-VTON,一种两阶段自回归框架,在VAR-VTON基础上解耦隐空间结构合成与像素空间细节恢复。核心思想是引入匹配感知重构器,建立第一阶段生成结果与源服装之间的密集对应关系,直接映射像素级细节。大量实验表明,STAR-VTON在效率与保真度间取得出色平衡:VAR-VTON推理速度至少比扩散模型快4倍,且不降低质量;像素空间重构器有效恢复细节,可作为即插即用模块提升现有方法性能。

原文摘要 · Abstract (English)

Virtual try-on (VTON) is a bi-conditional image generation problem that requires not only accurate person preservation but also faithful garment deformation and detail synthesis. Diffusion-based VTON methods can jointly model these factors in a compressed latent space, but suffer from high-frequency detail loss due to inherent latent compression, even with costly multi-step denoising. Recent visual autoregressive (VAR) models offer a promising alternative for high-quality generation with faster inference, yet remain unexplored for VTON due to the lack of effective bi-conditioning mechanisms. To bridge this gap, we first introduce VAR-VTON, a VAR-based VTON model that incorporates garment conditioning and structural guidance for efficient latent-space VTON. Despite its efficacy, latent-space generation still struggles to preserve fine-grained garment details. We argue that different VTON sub-tasks should be addressed in different representation spaces: structural synthesis such as garment warping and person layout is suited to the latent space, whereas fine-grained detail recovery should be tackled in the pixel space. Motivated by this insight, we further propose STAR-VTON, a Two-Stage AutoRegressive framework that builds upon VAR-VTON by decoupling latent-space structural synthesis from pixel-space detail recovery. Our idea is to resort to a matching-informed refiner to establish dense correspondences between the stage-one generation and the source garment to directly map fine-grained pixel-space details. Extensive experiments show that STAR-VTON achieves an impressive efficiency-fidelity trade-off: VAR-VTON runs at least $4\times$ faster than diffusion-based counterparts without degrading quality, and the pixel-space refiner effectively restores fine details and acts as a plug-and-play module that can benefit existing VTON approaches.

虚拟试穿自回归生成细节恢复图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。