arXiv:2507.05790cs.CV2025-07

用文字指令自动完成换装和局部修改的虚拟试穿助手

TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model

  • 基于大语言模型理解文本指令,动态切换不同处理流程
  • 无需手动画掩码即可实现全自动局部重绘
  • 支持全穿搭更换与局部编辑,灵活性更强

近年来,虚拟试穿技术取得显著进展。本文提出TalkFashion,一种仅通过文本指令即可实现多功能虚拟试穿的智能助手,包括整套服装更换和局部编辑。以往方法主要依赖端到端网络执行单一试穿任务,缺乏灵活性。我们利用大语言模型强大的理解能力分析用户指令,决定执行的任务并激活相应处理流程。此外,引入基于指令的局部重绘模型,无需用户手动提供掩码。借助多模态模型,该方法实现了全自动局部编辑,提升了编辑任务的灵活性。实验结果表明,相比现有方法,其语义一致性与视觉质量更优。

原文摘要 · Abstract (English)

Virtual try-on has made significant progress in recent years. This paper addresses how to achieve multifunctional virtual try-on guided solely by text instructions, including full outfit change and local editing. Previous methods primarily relied on end-to-end networks to perform single try-on tasks, lacking versatility and flexibility. We propose TalkFashion, an intelligent try-on assistant that leverages the powerful comprehension capabilities of large language models to analyze user instructions and determine which task to execute, thereby activating different processing pipelines accordingly. Additionally, we introduce an instruction-based local repainting model that eliminates the need for users to manually provide masks. With the help of multi-modal models, this approach achieves fully automated local editings, enhancing the flexibility of editing tasks. The experimental results demonstrate better semantic consistency and visual quality compared to the current methods.

虚拟试穿多模态大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。