用自然语言控制虚拟试穿,自动生成遮罩解决复杂款式难题
InstructVTON: Optimal Auto-Masking and Natural-Language-Guided Interactive Style Control for Inpainting-Based Virtual Try-On
- 通过视觉语言模型自动根据指令生成试穿遮罩
- 支持长袖卷起等复杂样式,突破传统遮罩限制
- 兼容现有模型,适合电商和时尚设计场景
我们提出 InstructVTON,一个遵循自然语言指令的交互式虚拟试穿系统,可对单件或多件服装进行精细复杂的风格控制。该系统将虚拟试穿建模为图像引导或图像条件的修复任务,采用计算高效且可扩展的方法。现有基于修复的虚拟试穿模型通常使用二值掩码控制生成布局,但生成理想掩码困难,需背景知识,可能依赖模型,且在某些情况下不可行(例如:给穿着长袖下摆的人试穿长袖卷袖,掩码必然覆盖整条袖子)。InstructVTON利用视觉语言模型(VLMs)和图像分割模型,基于用户提供的图像和自由文本风格指令自动生成掩码。该系统简化了用户体验,无需精确绘制掩码,并自动化执行多轮生成以实现仅靠掩码方法无法完成的试穿场景。我们证明 InstructVTON 可与现有虚拟试穿模型互操作,在保持风格控制的同时达到当前最优效果。
原文摘要 · Abstract (English)
We present InstructVTON, an instruction-following interactive virtual try-on system that allows fine-grained and complex styling control of the resulting generation, guided by natural language, on single or multiple garments. A computationally efficient and scalable formulation of virtual try-on formulates the problem as an image-guided or image-conditioned inpainting task. These inpainting-based virtual try-on models commonly use a binary mask to control the generation layout. Producing a mask that yields desirable result is difficult, requires background knowledge, might be model dependent, and in some cases impossible with the masking-based approach (e.g. trying on a long-sleeve shirt with "sleeves rolled up" styling on a person wearing long-sleeve shirt with sleeves down, where the mask will necessarily cover the entire sleeve). InstructVTON leverages Vision Language Models (VLMs) and image segmentation models for automated binary mask generation. These masks are generated based on user-provided images and free-text style instructions. InstructVTON simplifies the end-user experience by removing the necessity of a precisely drawn mask, and by automating execution of multiple rounds of image generation for try-on scenarios that cannot be achieved with masking-based virtual try-on models alone. We show that InstructVTON is interoperable with existing virtual try-on models to achieve state-of-the-art results with styling control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。