arXiv:2607.14807cs.CV2026-07

无需分割图,可同时试穿多件衣服并保留细节纹理。

TAMF-VTON: Texture-Aware Mask-Free Virtual Try-On via High-Fidelity Image Synthesis

论文配图:TAMF-VTON: Texture-Aware Mask-Free Virtual Try-On via High-Fidelity Image Synthesis
图 1 · 摘自论文原文
  • 用轻量级专家混合模型实现高效微调,不损失基础模型能力。
  • 通过频域监督优化高频一致性,精准还原衣物纹理细节。
  • 支持任意搭配多件服装,适合电商场景的快速部署。

基于扩散模型的虚拟试衣方法受限于依赖分割掩码、细粒度纹理保留不足以及对多服装组合支持有限,难以在真实电商场景中应用。本文提出TAMF-VTON,一种无需掩码、具备纹理感知能力的框架,在无约束条件下实现高保真图像合成。该方法在推理时无需人工解析或修复掩码,支持多种风格、类别和数量的服装同时试穿,完整保留人体结构与复杂纹理。其统一生成流程包含三个核心组件:(1)轻量级专家混合(MoE)适配方案,实现高效微调且不损害模型通用编辑能力;(2)频域监督机制,显式优化高频谱一致性以保持高保真纹理;(3)基于自适应修复策略的数据清洗管道,模拟逆向试衣过程生成高质量训练样本。大量实验表明,本方法在定量指标与视觉质量上均优于现有最先进方法。经量化优化后,模型在NVIDIA RTX 4090上每张图像推理时间低于15秒,为数字时尚场景中的规模化商用提供了可行方案。项目地址:https://www.style3d.ai/ai-photoshoot/virtual-clothing-try-on。

原文摘要 · Abstract (English)

Recent diffusion-based virtual try-on (VTON) methods remain limited by their reliance on segmentation masks, insufficient preservation of fine-grained textures, and limited support for arbitrary multi-garment compositions. Consequently, existing approaches still face significant challenges in real-world e-commerce deployment. We present TAMF-VTON, a texture-aware, mask-free framework that enables high-fidelity image synthesis under practical unconstrained conditions. Our method requires no human parsing or inpainting masks at inference time and supports diverse garment styles, categories, and quantities, enabling the simultaneous transfer of multiple items while preserving body structure and intricate texture details. This is achieved through a unified generative pipeline with three key components: (1) a lightweight Mixture-of-Experts (MoE) adaptation scheme that enables efficient fine-tuning without compromising the base model's general editing capabilities; (2) a frequency-domain supervision mechanism that explicitly optimizes high-frequency spectral consistency to preserve high-fidelity textures; and (3) a robust data curation pipeline employing an adaptive inpainting strategy to simulate the inverse VTON process for high-quality training pair generation. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods in both quantitative metrics and perceptual quality. Optimized for efficiency, the model achieves inference in under 15 seconds per image on an NVIDIA RTX 4090 with INT4 quantization. By combining mask-free operation, flexible multi-garment composition, faithful texture preservation, and efficient inference on consumer hardware, TAMF-VTON demonstrates a commercially viable solution for scalable deployment in real-world digital fashion scenarios. The project is available at https://www.style3d.ai/ai-photoshoot/virtual-clothing-try-on.

虚拟试衣扩散模型纹理保真高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。