arXiv:2605.14874cs.CV2026-05

提出新框架LPH-VTON,解决虚拟试穿中结构与纹理的矛盾。

LPH-VTON: Resolving the Structure-Texture Dilemma of Virtual Try-On via Latent Process Handover

论文配图:LPH-VTON: Resolving the Structure-Texture Dilemma of Virtual Try-On via Latent Process Handover
图 1 · 摘自论文原文
  • 早期用结构模型建骨架,后期交由纹理模型精修。
  • 在VITON-HD上同时提升外观真实感和身体对齐度。
  • 适合关注生成质量与形体一致性的视觉生成研究者。

虚拟试穿(VTON)旨在生成与人体姿态精准对齐的逼真服装图像。现有基于扩散的方法在结构完整性和纹理保真度之间面临根本性权衡。本文将此问题归因于现有架构内在互补的归纳偏置:依赖空间约束的模型更利于几何对齐但常抑制纹理细节,而依赖无约束生成先验的模型擅长渲染丰富细节却易出现结构漂移。为此,我们提出LPH-VTON,一种协同框架,在单一连续去噪过程中化解该矛盾。该方法分阶段生成:前期由结构偏向模型建立几何一致的潜在骨架,随后移交控制权给纹理偏向模型进行高保真细节渲染。大量实验验证了该方法的有效性。模型在标准数据集VITON-HD上实现了更优的帕累托平衡,显著提升了感知真实感,同时保持高度竞争性的结构对齐性能,证明了时间上架构解耦的有效性。

原文摘要 · Abstract (English)

Virtual Try-On (VTON) aims to synthesize photorealistic images of garments precisely aligned with a person's body and pose. Current diffusion-based methods, however, face a fundamental trade-off between structural integrity and textural fidelity. In this paper, we formalize this challenge as a consequence of complementary inductive biases inherent in prevailing architectures: models heavily reliant on spatial constraints naturally favor geometric alignment but often suppress textures, whereas models dominated by unconstrained generative priors excel at vibrant detail rendering but are prone to structural drift. Based on this diagnosis, we propose LPH-VTON, a new synergistic framework that resolves this tension within a single, continuous denoising process. LPH-VTON strategically decomposes the generation, leveraging a structure-biased model to establish a geometrically consistent latent scaffold in the early stages, before handing over control to a texture-biased model for high-fidelity detail rendering. Extensive experiments validate our approach. Our model achieves a superior Pareto-optimal balance, establishing new benchmarks in perceptual faithfulness while maintaining highly competitive structural alignment across the standard dataset VITON-HD, proving the efficacy of temporal architectural decoupling.

虚拟试穿扩散模型结构纹理协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。