arXiv:2512.07126cs.CV2025-12

无需训练即可提升虚拟试衣的服装细节匹配度

Training-free Clothing Region of Interest Self-correction for Virtual Try-On

  • 用能量函数约束生成过程中的注意力,聚焦服装区域
  • 在多个指标上超越现有方法,最高提升12.3%
  • 适合需要高精度服装还原的应用场景

虚拟试衣(VTON)旨在将目标服装合成到特定人物身上,保持服装细节的同时使人物其他部分不变。现有方法在图案、纹理和边界上与目标服装存在差异。为此,我们提出通过能量函数约束生成过程中提取的注意力图,使每一步生成时注意力更聚焦于服装感兴趣区域,从而提升生成结果与目标服装细节的一致性。此外,针对现有评估指标仅关注图像真实感而忽略与目标元素对齐的问题,我们设计了新的评估指标虚拟试衣初始距离(VTID),以实现更全面的评价。在VITON-HD和DressCode数据集上,我们的方法在传统指标LPIPS、FID、KID及新指标VTID上分别优于先前最先进方法1.4%、2.3%、12.3%和5.8%。进一步将生成数据用于下游服装变化重识别任务,在LTCC、PRCC、VC-Clothes数据集的Rank-1指标上分别提升2.5%、1.1%和1.6%。代码已开源:https://github.com/MrWhiteSmall/CSC-VTON.git。

原文摘要 · Abstract (English)

VTON (Virtual Try-ON) aims at synthesizing the target clothing on a certain person, preserving the details of the target clothing while keeping the rest of the person unchanged. Existing methods suffer from the discrepancies between the generated clothing results and the target ones, in terms of the patterns, textures and boundaries. Therefore, we propose to use an energy function to impose constraints on the attention map extracted through the generation process. Thus, at each generation step, the attention can be more focused on the clothing region of interest, thereby influencing the generation results to be more consistent with the target clothing details. Furthermore, to address the limitation that existing evaluation metrics concentrate solely on image realism and overlook the alignment with target elements, we design a new metric, Virtual Try-on Inception Distance (VTID), to bridge this gap and ensure a more comprehensive assessment. On the VITON-HD and DressCode datasets, our approach has outperformed the previous state-of-the-art (SOTA) methods by 1.4%, 2.3%, 12.3%, and 5.8% in the traditional metrics of LPIPS, FID, KID, and the new VTID metrics, respectively. Additionally, by applying the generated data to downstream Clothing-Change Re-identification (CC-Reid) methods, we have achieved performance improvements of 2.5%, 1.1%, and 1.6% on the LTCC, PRCC, VC-Clothes datasets in the metrics of Rank-1. The code of our method is public at https://github.com/MrWhiteSmall/CSC-VTON.git.

虚拟试衣注意力控制图像生成评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。