arXiv:2412.11513cs.CV2024-12被引 8

用双提取器和融合块提升服装还原质量,保留原服装特征。

IGR: Improving Diffusion Model for Garment Restoration from Person Image

  • 双提取器分别捕捉服装低层特征与高层语义。
  • 融合块结合自注意力与交叉注意力,对齐人物与服装。
  • 粗到精训练策略增强生成服装的真实感与细节。

服装还原是虚拟试穿的逆任务,旨在从人物图像中恢复标准服装,需精确捕捉服装细节。现有方法常无法保持服装身份或依赖复杂流程。为此,我们提出一种改进的扩散模型,用于恢复真实服装。该方法采用两个服装提取器,分别从人物图像中独立提取低层特征与高层语义。利用预训练的潜在扩散模型,这些特征通过服装融合块融入去噪过程,融合块结合自注意力与交叉注意力层,使重建服装与人物图像对齐。此外,引入粗到精训练策略,提升生成服装的保真度与真实性。实验表明,本模型能有效保持服装身份,在复杂服装或遮挡场景下仍生成高质量还原结果。

原文摘要 · Abstract (English)

Garment restoration, the inverse of virtual try-on task, focuses on restoring standard garment from a person image, requiring accurate capture of garment details. However, existing methods often fail to preserve the identity of the garment or rely on complex processes. To address these limitations, we propose an improved diffusion model for restoring authentic garments. Our approach employs two garment extractors to independently capture low-level features and high-level semantics from the person image. Leveraging a pretrained latent diffusion model, these features are integrated into the denoising process through garment fusion blocks, which combine self-attention and cross-attention layers to align the restored garment with the person image. Furthermore, a coarse-to-fine training strategy is introduced to enhance the fidelity and authenticity of the generated garments. Experimental results demonstrate that our model effectively preserves garment identity and generates high-quality restorations, even in challenging scenarios such as complex garments or those with occlusions.

服装还原扩散模型图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。