统一生成与虚拟试穿,实现服装设计全流程可控合成
VersaVogue: Visual Expert Orchestration and Preference Alignment for Unified Fashion Synthesis
- 用专家路由注意力动态分离纹理、形状、颜色等属性
- 无需人工标注即可构建偏好数据,提升真实感与可控性
- 适合需要精细控制的时尚生成场景,如品牌设计与电商展示
扩散模型推动了时尚图像生成的显著进展,但现有方法通常将服装生成与虚拟试穿视为独立问题,限制了实际应用中的灵活性。此外,在多源异构条件下的时尚图像合成仍具挑战,因现有方法多依赖简单的特征拼接或静态层注入,易导致属性纠缠与语义干扰。为此,我们提出 VersaVogue,一个统一的多条件可控时尚合成框架,同时支持服装生成与虚拟试穿,对应时尚生命周期的设计与展示阶段。具体地,引入特质路由注意力(TA)模块,利用混合专家机制动态将条件特征路由至最匹配的专家与生成层,实现纹理、形状、颜色等视觉属性的解耦注入。为进一步提升真实感与可控性,构建无需人工标注或特定任务奖励模型的自动化多视角偏好优化(MPO)流程。通过内容保真度、文本对齐性与感知质量评估器构建可靠偏好对,并采用直接偏好优化(DPO)进行模型训练。在服装生成与虚拟试穿基准上的大量实验表明,VersaVogue在视觉保真度、语义一致性与细粒度可控性方面均持续优于现有方法。
原文摘要 · Abstract (English)
Diffusion models have driven remarkable advancements in fashion image generation, yet prior works usually treat garment generation and virtual dressing as separate problems, limiting their flexibility in real-world fashion workflows. Moreover, fashion image synthesis under multi-source heterogeneous conditions remains challenging, as existing methods typically rely on simple feature concatenation or static layer-wise injection, which often causes attribute entanglement and semantic interference. To address these issues, we propose VersaVogue, a unified framework for multi-condition controllable fashion synthesis that jointly supports garment generation and virtual dressing, corresponding to the design and showcase stages of the fashion lifecycle. Specifically, we introduce a trait-routing attention (TA) module that leverages a mixture-of-experts mechanism to dynamically route condition features to the most compatible experts and generative layers, enabling disentangled injection of visual attributes such as texture, shape, and color. To further improve realism and controllability, we develop an automated multi-perspective preference optimization (MPO) pipeline that constructs preference data without human annotation or task-specific reward models. By combining evaluators of content fidelity, textual alignment, and perceptual quality, MPO identifies reliable preference pairs, which are then used to optimize the model via direct preference optimization (DPO). Extensive experiments on both garment generation and virtual dressing benchmarks demonstrate that VersaVogue consistently outperforms existing methods in visual fidelity, semantic consistency, and fine-grained controllability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。