arXiv:2606.27773cs.CV2026-06

通过模态感知流匹配,实现高保真虚拟试衣,精准对齐文本与服装外观。

ModaFlow: Modality-Aware Flow Matching for High-Fidelity Virtual Try-On

论文配图:ModaFlow: Modality-Aware Flow Matching for High-Fidelity Virtual Try-On
图 1 · 摘自论文原文
  • 分模态引导:视觉嵌入提供稳定结构,文本嵌入用自适应无分类器指导。
  • 新损失函数提升流场方向一致性和感知真实感,降低FID达30%以上。
  • 随机掩码策略增强泛化能力,适合无配对数据下的鲁棒试衣应用。

基于图像的虚拟试衣在电商和增强现实领域日益重要,但现有方法难以同时保持服装细节语义并适应大形变下多样的人体几何。我们提出ModaFlow,一种基于模态感知流匹配的高保真虚拟试衣框架,能精确对齐文本描述与服装外观。不同于以往统一处理多模态条件的方法,ModaFlow引入分模态引导机制:由预训练图像提示适配器提取的视觉服装嵌入提供确定性、持续性的结构引导;而由服装描述生成的文本嵌入则通过无分类器引导(CFG)结合自适应缩放与零初始化速度进行控制。为进一步提升流场精度,我们提出余弦相似度和感知流判别两项正则化损失,联合优化速度场的方向一致性与感知真实性。此外,采用掩码操作策略,在训练中随机采样框掩码、透明掩码和松弛掩码,模拟多样遮挡场景,使模型在仅提供框掩码的无配对设置下仍具鲁棒性。实验表明,ModaFlow在定性和定量评估中均达到最优表现,在配对与无配对基准上分别将FID降低约30%和20%。

原文摘要 · Abstract (English)

Image-based virtual try-on has emerged as a compelling task in e-commerce and augmented reality, yet existing methods struggle to simultaneously preserve fine garment semantics and adapt to diverse person body geometries under large clothing-body deformations. We present ModaFlow, a modality-aware flow-matching based framework for high-fidelity virtual try-on that achieves precise alignment between textual descriptions and garment appearance. Unlike prior methods that treat multimodal conditions uniformly, ModaFlow introduces a modality-aware guidance scheme: visual garment embeddings extracted by a pretrained image prompt adapter provide deterministic, persistent structural guidance, while textual embeddings generated from garment descriptions are controlled via classifier-free guidance (CFG) with adaptive scaling and zero-initialized velocity. To further enhance flow field accuracy, we propose two regularization losses, cosine similarity and perceptual flow discrimination, that jointly improve directional consistency and perceptual realism of the velocity field. Additionally, a mask manipulation strategy stochastically samples among box, transparent, and relaxed masks during training, simulating diverse occlusion scenarios and enabling robust inference under unpaired settings where only a box mask is available. Experiments show that ModaFlow achieves state-of-the-art results in both qualitative and quantitative evaluations, reducing FID by approximately 30% on paired and 20% on unpaired benchmarks.

虚拟试衣流匹配多模态高保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。