让虚拟试衣更智能:可自定义模特姿态与外观的服装试穿系统
Neural Clothing Tryer: Customized Virtual Try-On via Semantic Enhancement and Controlling Diffusion Model
- 用语义增强与控制扩散模型,精准保留服装细节
- 支持自由调整模特姿态、表情和个性特征
- 适合电商试衣、数字人设计等场景使用
本文提出一种新型定制化虚拟试衣(Cu-VTON)任务,允许将指定服装叠加到可自定义外观、姿态及附加属性的数字模特上。相比传统虚拟试衣,该方法让用户能按个人偏好定制数字形象,显著提升试衣体验的灵活性与沉浸感。为此,我们提出神经服装试穿框架(NCT),利用具备语义增强与控制模块的先进扩散模型,以更好保持服装的语义特征与纹理细节,同时实现模特姿态与外观的灵活编辑。具体地,NCT引入语义增强模块,通过视觉-语言编码器学习跨模态对齐特征,并作为条件输入扩散模型以强化服装语义保留;进一步设计语义控制模块,接收服装图像、定制姿态图及语义描述,实现服装细节保持的同时,灵活编辑模特姿态、表情与多种属性。在公开基准数据集上的大量实验表明,所提NCT框架性能显著优于现有方法。
原文摘要 · Abstract (English)
This work aims to address a novel Customized Virtual Try-ON (Cu-VTON) task, enabling the superimposition of a specified garment onto a model that can be customized in terms of appearance, posture, and additional attributes. Compared with traditional VTON task, it enables users to tailor digital avatars to their individual preferences, thereby enhancing the virtual fitting experience with greater flexibility and engagement. To address this task, we introduce a Neural Clothing Tryer (NCT) framework, which exploits the advanced diffusion models equipped with semantic enhancement and controlling modules to better preserve semantic characterization and textural details of the garment and meanwhile facilitating the flexible editing of the model's postures and appearances. Specifically, NCT introduces a semantic-enhanced module to take semantic descriptions of garments and utilizes a visual-language encoder to learn aligned features across modalities. The aligned features are served as condition input to the diffusion model to enhance the preservation of the garment's semantics. Then, a semantic controlling module is designed to take the garment image, tailored posture image, and semantic description as input to maintain garment details while simultaneously editing model postures, expressions, and various attributes. Extensive experiments on the open available benchmark demonstrate the superior performance of the proposed NCT framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。