arXiv:2505.19638cs.CV2025-05被引 3

解决虚拟试衣中姿态变化时的形变与语义不一致问题,提升图像真实感。

HF-VTON: High-Fidelity Virtual Try-On via Consistent Geometric and Semantic Alignment

  • 分三模块对齐姿态、增强语义表征、融合多模态先验生成细节
  • 在多姿态下保持结构、纹理和图案一致性,显著提升视觉保真度
  • 适用于时尚电商高精度虚拟试衣,适合关注细节还原的研究者

虚拟试衣技术在时尚零售领域日益重要,可生成适配目标人体模型的高保真服装图像。现有方法在不同姿态下仍面临显著挑战:几何失真导致空间不一致,服装结构与纹理错位引发语义不一致,细粒度细节丢失降低视觉保真度。为此,本文提出HF-VTON框架,包含三个核心模块:(1) 风格保持的变形对齐模块(APWAM),实现服装到人体姿态的精准对齐,缓解几何形变并保证空间一致性;(2) 语义表征与理解模块(SRCM),通过捕捉细粒度服装属性及多姿态数据,强化语义表达,维持结构、纹理与图案的一致性;(3) 多模态先验引导的外观生成模块(MPAGM),融合多模态特征与预训练模型先验知识,优化外观生成,确保语义与几何双重一致性。此外,为克服现有基准数据限制,我们构建了SAMP-VTONS数据集,包含多姿态配对与丰富文本标注,支持更全面评估。实验表明,HF-VTON在VITON-HD与SAMP-VTONS上均优于当前最优方法,在视觉保真度、语义一致性和细节保留方面表现优异。

原文摘要 · Abstract (English)

Virtual try-on technology has become increasingly important in the fashion and retail industries, enabling the generation of high-fidelity garment images that adapt seamlessly to target human models. While existing methods have achieved notable progress, they still face significant challenges in maintaining consistency across different poses. Specifically, geometric distortions lead to a lack of spatial consistency, mismatches in garment structure and texture across poses result in semantic inconsistency, and the loss or distortion of fine-grained details diminishes visual fidelity. To address these challenges, we propose HF-VTON, a novel framework that ensures high-fidelity virtual try-on performance across diverse poses. HF-VTON consists of three key modules: (1) the Appearance-Preserving Warp Alignment Module (APWAM), which aligns garments to human poses, addressing geometric deformations and ensuring spatial consistency; (2) the Semantic Representation and Comprehension Module (SRCM), which captures fine-grained garment attributes and multi-pose data to enhance semantic representation, maintaining structural, textural, and pattern consistency; and (3) the Multimodal Prior-Guided Appearance Generation Module (MPAGM), which integrates multimodal features and prior knowledge from pre-trained models to optimize appearance generation, ensuring both semantic and geometric consistency. Additionally, to overcome data limitations in existing benchmarks, we introduce the SAMP-VTONS dataset, featuring multi-pose pairs and rich textual annotations for a more comprehensive evaluation. Experimental results demonstrate that HF-VTON outperforms state-of-the-art methods on both VITON-HD and SAMP-VTONS, excelling in visual fidelity, semantic consistency, and detail preservation.

虚拟试衣姿态对齐语义一致性图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。