提出CORAL框架,提升虚拟试衣中衣物与人体的对齐精度
CORAL: Correspondence Alignment for Improved Virtual Try-On
- 通过显式对齐查询-键匹配,增强扩散变换器中的衣物与人体对应关系
- 在未配对设置下,显著改善全局形态迁移和局部细节保留
- 适合关注虚拟试衣细节还原的研究者与工业应用开发者
现有虚拟试衣方法在未配对场景下难以保持精细衣物细节,因缺乏对人体-衣物对应关系的显式约束,且无法解释扩散变换器(DiT)中对应关系如何形成。本文首次分析基于DiT架构的完整3D注意力机制,揭示人体-衣物对应关系依赖于3D注意力中精确的人体-衣物查询-键匹配。据此,提出CORAL框架,通过两个互补模块实现显式对齐:对应关系蒸馏损失将可靠匹配与人体-衣物注意力对齐,熵最小化损失使注意力分布更集中。同时提出基于视觉语言模型(VLM)的评估协议,更贴近人类偏好。实验表明,CORAL在多个指标上持续优于基线,有效提升全局形态转移与局部细节保留能力。大量消融实验证明设计合理性。
原文摘要 · Abstract (English)
Existing methods for Virtual Try-On (VTON) often struggle to preserve fine garment details, especially in unpaired settings where accurate person-garment correspondence is required. These methods do not explicitly enforce person-garment alignment and fail to explain how correspondence emerges within Diffusion Transformers (DiTs). In this paper, we first analyze full 3D attention in DiT-based architecture and reveal that the person-garment correspondence critically depends on precise person-garment query-key matching within the full 3D attention. Building on this insight, we then introduce CORrespondence ALignment (CORAL), a DiT-based framework that explicitly aligns query-key matching with robust external correspondences. CORAL integrates two complementary components: a correspondence distillation loss that aligns reliable matches with person-garment attention, and an entropy minimization loss that sharpens the attention distribution. We further propose a VLM-based evaluation protocol to better reflect human preference. CORAL consistently improves over the baseline, enhancing both global shape transfer and local detail preservation. Extensive ablations validate our design choices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。