通过特征对齐提升虚拟试穿图像细节真实度
AlignVTOFF: Texture-Spatial Feature Alignment for High-Fidelity Virtual Try-Off
- 设计并行U-Net结构,融合参考图与纹理空间特征
- 在多个数据集上优于现有方法,显著保留高频率纹理细节
- 适合需要高精度服装生成的应用场景
虚拟试穿(VTOFF)是一项具有挑战性的多模态图像生成任务,旨在合成复杂几何形变和丰富高频纹理下的高质量平铺服装图像。现有方法多依赖轻量模块进行快速特征提取,难以保持结构化图案和细微细节,导致生成过程中纹理衰减。为此,我们提出AlignVTOFF,一种基于参考U-Net与纹理-空间特征对齐(TSFA)的新型并行U-Net框架。参考U-Net实现多尺度特征提取,增强几何保真度,有效建模形变并保留复杂结构模式。TSFA通过混合注意力设计,将参考服装特征注入冻结的去噪U-Net,包含可训练交叉注意力与冻结自注意力模块,显式对齐纹理与空间线索,缓解去噪过程中的高频信息丢失。大量实验表明,AlignVTOFF在多种设置下持续优于当前最优方法,在结构真实性和高频细节保真度上均有显著提升。
原文摘要 · Abstract (English)
Virtual Try-Off (VTOFF) is a challenging multimodal image generation task that aims to synthesize high-fidelity flat-lay garments under complex geometric deformation and rich high-frequency textures. Existing methods often rely on lightweight modules for fast feature extraction, which struggles to preserve structured patterns and fine-grained details, leading to texture attenuation during generation.To address these issues, we propose AlignVTOFF, a novel parallel U-Net framework built upon a Reference U-Net and Texture-Spatial Feature Alignment (TSFA). The Reference U-Net performs multi-scale feature extraction and enhances geometric fidelity, enabling robust modeling of deformation while retaining complex structured patterns. TSFA then injects the reference garment features into a frozen denoising U-Net via a hybrid attention design, consisting of a trainable cross-attention module and a frozen self-attention module. This design explicitly aligns texture and spatial cues and alleviates the loss of high-frequency information during the denoising process.Extensive experiments across multiple settings demonstrate that AlignVTOFF consistently outperforms state-of-the-art methods, producing flat-lay garment results with improved structural realism and high-frequency detail fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。