arXiv:2603.09236cs.CVcs.AI2026-03

用扩散模型连接人体观察与平铺服装,提升虚拟试穿的还原精度

BridgeDiff: Bridging Human Observations and Flat-Garment Synthesis for Virtual Try-Off

  • 通过双模块设计,融合全局服装特征与结构先验
  • 在标准数据集上实现当前最优的平铺服装重建质量
  • 适合需要高保真服装建模的电商与服装设计场景

虚拟试穿(VTOFF)旨在从穿在人身上的图像中恢复标准化的平铺服装表示,以支持统一展示和下游虚拟试穿任务。现有方法常将VTOFF视为仅依赖局部掩码或纯文本提示的直接图像转换,忽略了人体穿着外观与平铺布局之间的差异,导致未观测区域补全不一致且服装结构不稳定。本文提出BridgeDiff,一种基于扩散模型的框架,通过两个互补组件显式桥接人体中心观察与平铺服装生成。首先,服装线索桥接模块(GCBM)构建捕捉全局外观与语义身份的服装线索表征,实现部分可见情况下的连续细节鲁棒推断。其次,平铺结构约束模块(FSCM)在特定去噪阶段通过平铺-约束注意力(FC-Attention)注入显式平铺结构先验,提升结构稳定性,超越仅依赖文本的条件控制。在标准VTOFF基准上的大量实验表明,BridgeDiff达到当前最优性能,生成更高质量的平铺服装重建结果,同时保留精细外观与结构完整性。

原文摘要 · Abstract (English)

Virtual try-off (VTOFF) aims to recover canonical flat-garment representations from images of dressed persons for standardized display and downstream virtual try-on. Prior methods often treat VTOFF as direct image translation driven by local masks or text-only prompts, overlooking the gap between on-body appearances and flat layouts. This gap frequently leads to inconsistent completion in unobserved regions and unstable garment structure. We propose BridgeDiff, a diffusion-based framework that explicitly bridges human-centric observations and flat-garment synthesis through two complementary components. First, the Garment Condition Bridge Module (GCBM) builds a garment-cue representation that captures global appearance and semantic identity, enabling robust inference of continuous details under partial visibility. Second, the Flat Structure Constraint Module (FSCM) injects explicit flat-garment structural priors via Flat-Constraint Attention (FC-Attention) at selected denoising stages, improving structural stability beyond text-only conditioning. Extensive experiments on standard VTOFF benchmarks show that BridgeDiff achieves state-of-the-art performance, producing higher-quality flat-garment reconstructions while preserving fine-grained appearance and structural integrity.

虚拟试穿扩散模型服装生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。