arXiv:2511.14031cs.CV2025-11AAAI

无需变形生成服装图像,精细控制模特外观。

FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance Customization

  • 不进行服装变形,直接外推分割出的服装区域。
  • 通过区域自适应注意力机制实现属性精准控制。
  • 适合电商场景中高保真服装展示生成。

以服装为中心的时尚图像生成旨在合成穿着指定服装的真实且可控制的人体模型,因其在电子商务中的应用前景而受到广泛关注。该任务的核心挑战在于:(1) 精确保留服装细节,(2) 实现对模型外观的细粒度控制。现有方法通常需在生成过程中进行服装变形,常导致纹理失真;同时因缺乏专门设计的机制,难以控制生成模型的细粒度属性。为此,我们提出FashionMAC,一种基于扩散模型的无变形框架,实现高质量、可控制的时尚展示图像生成。其核心思想是消除服装变形需求,直接外推从着装人体中分割出的服装区域,从而忠实保留复杂服装细节。此外,我们提出一种新型区域自适应解耦注意力(RADA)机制,并结合链式掩码注入策略,实现对合成人体模型的细粒度外观控制。具体而言,RADA 自适应预测每项细粒度文本属性的生成区域,并通过链式掩码注入策略强制文本属性聚焦于预测区域,显著提升视觉保真度与可控性。大量实验验证了本框架相比现有最先进方法的优越性能。

原文摘要 · Abstract (English)

Garment-centric fashion image generation aims to synthesize realistic and controllable human models dressing a given garment, which has attracted growing interest due to its practical applications in e-commerce. The key challenges of the task lie in two aspects: (1) faithfully preserving the garment details, and (2) gaining fine-grained controllability over the model's appearance. Existing methods typically require performing garment deformation in the generation process, which often leads to garment texture distortions. Also, they fail to control the fine-grained attributes of the generated models, due to the lack of specifically designed mechanisms. To address these issues, we propose FashionMAC, a novel diffusion-based deformation-free framework that achieves high-quality and controllable fashion showcase image generation. The core idea of our framework is to eliminate the need for performing garment deformation and directly outpaint the garment segmented from a dressed person, which enables faithful preservation of the intricate garment details. Moreover, we propose a novel region-adaptive decoupled attention (RADA) mechanism along with a chained mask injection strategy to achieve fine-grained appearance controllability over the synthesized human models. Specifically, RADA adaptively predicts the generated regions for each fine-grained text attribute and enforces the text attribute to focus on the predicted regions by a chained mask injection strategy, significantly enhancing the visual fidelity and the controllability. Extensive experiments validate the superior performance of our framework compared to existing state-of-the-art methods.

服装生成扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。