arXiv:2607.09362cs.CVcs.AI2026-07

让用户像剪裁衣服一样精细控制虚拟试穿的版型与穿搭方式。

CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation

论文配图:CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation
图 1 · 摘自论文原文
  • 通过视觉实例提示分割,精准定位服装在人体上的具体位置。
  • 生成结果严格遵循用户指定的版型布局,且服装还原度媲美顶级商业系统。
  • 适合需要高度定制化试穿效果的设计、电商或内容创作场景。

虚拟试穿(VTO)在将衣物真实地转移到目标人物上已取得显著进展。然而,大多数系统对用户如何穿着衣物的控制能力较弱——如衣物大小(宽松或合身)、风格(如塞入或未塞入、敞开或闭合)以及在身体上的空间位置。本文提出两项互补贡献:首先,定义并解决视觉实例提示分割(VIP-SAM),即给定一件服装的平铺图,将其在一张人物穿着照片中特定实例精确分割出来,这是一个实例级任务,不同于通常研究的类别级分割;其次,提出可控虚拟试穿框架CtrlVTON,将试穿重构为图像编辑问题,并引入分割掩码作为像素级控制,以实现对衣物版型、风格及空间布局的精确调控。VIP-SAM和CtrlVTON分别在其任务上达到当前最优表现。特别地,CtrlVTON生成的图像比最强的专有编辑系统更忠实地遵循用户提供的布局,同时在服装保真度上与之相当。

原文摘要 · Abstract (English)

Virtual try-on (VTO) has made significant progress in realistically transferring garments onto a target person. Yet most systems give the user little control over how a garment should be worn -- its size (loose or fitted), style (e.g., tucked in or untucked, open or closed), and spatial placement on the body. We address this gap with two complementary contributions. First, we define and solve Visual-Instance-Prompt Segmentation via VIP-SAM: given a flatlay image of a garment, segment that specific instance in a photograph of a person wearing it. This is an instance-level task, distinct from the typically studied category-level segmentation. Second, we introduce CtrlVTON, a controllable VTO framework that recasts try-on as an image editing problem and adds segmentation masks as pixel-level control over garment layout, including style, size, and spatial placement on the body. VIP-SAM and CtrlVTON each achieve state-of-the-art results on their respective tasks. In particular, CtrlVTON generates images that follow user-provided layouts far more faithfully than the strongest proprietary editing systems while matching them on garment fidelity.

虚拟试穿图像编辑实例分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。