用传统特征点匹配提升虚拟试穿的细节精度
SIFT-VTON: Geometric Correspondence Supervision on Cross-Attention for Virtual Try-On

- 通过SIFT关键点匹配提供显式几何引导
- 在VITON-HD上未配对指标显著提升
- 适合关注服装纹理与图案对齐的研究者
基于扩散模型的虚拟试穿方法通过交叉注意力机制将服装特征转移到目标身体区域,但依赖隐式空间对应关系,难以保留文字和图案等细粒度信息。本文提出SIFT-VTON,利用SIFT关键点匹配为扩散模型提供显式几何指导。方法对服装与人体图像间的SIFT匹配结果进行领域特定过滤,并将其转换为空间概率分布,在训练中监督交叉注意力层。该显式引导促使模型学习精确的空间对齐,注意力集中于几何一致的服装区域。在VITON-HD数据集上的实验表明,未配对指标显著提升,同时保持了具有竞争力的配对重建性能。定性对比显示文本清晰度和图案对齐效果更优。注意力可视化证实本方法产生聚焦于关键服装细节的锐利注意力。结果表明,经典几何对应方法可有效增强现代扩散模型在条件生成任务中的表现。
原文摘要 · Abstract (English)
Diffusion-based virtual try-on methods achieve photorealistic synthesis through cross-attention mechanisms that transfer garment features to target body regions. However, these approaches rely on implicit learning of spatial correspondences, struggling to preserve fine details such as text and illustrations. We propose a novel approach, which we call SIFT-VTON, that utilizes SIFT keypoint matching to provide explicit geometric guidance for diffusion-based virtual try-on. Our method applies domain-specific filtering to SIFT keypoint matches between garment and person images, then converts these correspondences into spatial probability distributions that supervise cross-attention layers during training. This explicit supervision guides the model to learn precise spatial alignment, concentrating attention on geometrically consistent garment regions. Experiments on the VITON-HD dataset demonstrate significant improvements on unpaired metrics while maintaining competitive paired reconstruction metrics. Qualitative comparisons show superior preservation of text clarity and pattern alignment. Attention visualizations confirm that our method produces sharply focused attention on relevant garment details. This work demonstrates that classical geometric correspondence methods can effectively enhance modern diffusion models for conditional synthesis tasks. The source code will be available at https://github.com/takesukeDS/SIFT-VTON.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。