arXiv:2509.20343cs.CV2025-09被引 4

不加额外模块,用姿态图实现高效虚拟试穿的姿势控制。

Efficient Encoder-Free Pose Conditioning and Pose Control for Virtual Try-On

  • 纯拼接架构下,通过空间拼接姿态图实现无参数姿势控制。
  • 姿态图拼接比骨骼结构更能保持姿势准确性与图像真实感。
  • 结合细粒度与框选掩码训练,支持多姿态下的灵活服装融合。

随着在线购物增长,虚拟试穿(VTON)技术需求激增,使用户能将商品图像叠加到自身照片上预览效果。姿势控制是实现有效VTON的关键挑战,需确保商品与身体精准对齐,并支持多样姿态以增强沉浸感。现有方法在姿态表示选择、无额外参数集成及姿态保真与灵活控制间平衡方面存在困难。本文基于一个不依赖外部编码器、控制网络或复杂注意力层的基线VTON模型,探索在纯拼接框架下融入姿态控制的方法:通过空间拼接姿态数据,比较使用姿态图与骨架的表现,且不引入任何新参数或模块。实验表明,姿态图拼接效果最佳,显著提升姿态保真度与输出真实性。此外,提出一种混合掩码训练策略,结合细粒度掩码与边界框掩码,使模型能在多种姿态和条件下灵活整合服装。

原文摘要 · Abstract (English)

As online shopping continues to grow, the demand for Virtual Try-On (VTON) technology has surged, allowing customers to visualize products on themselves by overlaying product images onto their own photos. An essential yet challenging condition for effective VTON is pose control, which ensures accurate alignment of products with the user's body while supporting diverse orientations for a more immersive experience. However, incorporating pose conditions into VTON models presents several challenges, including selecting the optimal pose representation, integrating poses without additional parameters, and balancing pose preservation with flexible pose control. In this work, we build upon a baseline VTON model that concatenates the reference image condition without external encoder, control network, or complex attention layers. We investigate methods to incorporate pose control into this pure concatenation paradigm by spatially concatenating pose data, comparing performance using pose maps and skeletons, without adding any additional parameters or module to the baseline model. Our experiments reveal that pose stitching with pose maps yields the best results, enhancing both pose preservation and output realism. Additionally, we introduce a mixed-mask training strategy using fine-grained and bounding box masks, allowing the model to support flexible product integration across varied poses and conditions.

虚拟试穿姿态控制无编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。