解决松身衣物虚拟试穿的错位与抖动问题
Real-Time Per-Garment Virtual Try-On with Temporal Consistency for Loose-Fitting Garments
- 用不变特征+辅助网络重建遮挡下的身体语义图
- 引入时序递归结构,实现真实帧间一致性
- 专为松身衣物设计,适合实时试穿系统
针对松身衣物虚拟试穿中因身体轮廓被遮挡导致的语义图失真及帧间抖动问题,本文提出两阶段解决方案。首先,通过提取服装无关特征并经辅助网络生成更鲁棒的语义图,提升在松身衣物下的对齐精度;其次,设计基于递归机制的合成框架,利用时序依赖性增强帧间一致性,同时保持实时性能。实验表明,该方法在图像质量与时间连贯性上均优于现有方法。消融实验验证了服装无关表示与递归架构的有效性。
原文摘要 · Abstract (English)
Per-garment virtual try-on methods collect garment-specific datasets and train networks tailored to each garment to achieve superior results. However, these approaches often struggle with loose-fitting garments due to two key limitations: (1) They rely on human body semantic maps to align garments with the body, but these maps become unreliable when body contours are obscured by loose-fitting garments, resulting in degraded outcomes; (2) They train garment synthesis networks on a per-frame basis without utilizing temporal information, leading to noticeable jittering artifacts. To address the first limitation, we propose a two-stage approach for robust semantic map estimation. First, we extract a garment-invariant representation from the raw input image. This representation is then passed through an auxiliary network to estimate the semantic map. This enhances the robustness of semantic map estimation under loose-fitting garments during garment-specific dataset generation. To address the second limitation, we introduce a recurrent garment synthesis framework that incorporates temporal dependencies to improve frame-to-frame coherence while maintaining real-time performance. We conducted qualitative and quantitative evaluations to demonstrate that our method outperforms existing approaches in both image quality and temporal coherence. Ablation studies further validate the effectiveness of the garment-invariant representation and the recurrent synthesis framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。