提出细粒度服装保真度评估框架,让虚拟试穿更精准
Beyond Global Realism: Virtual Try-On Evaluation and Optimization with Dimension-wise Garment Fidelity Assessment

- 将服装一致性拆解为7个可解释维度,逐项评估
- 在50K弱监督+10K精标数据上训练,平衡标签偏差
- 可嵌入强化学习优化生成模型,自动聚焦薄弱环节
虚拟试穿不仅需要图像真实,还需忠实保留服装特征。现有指标如PSNR、SSIM、KID和FID难以衡量生成与参考服装的一致性,尤其在捕捉多维服装保真度方面表现不足。为此,我们提出DAT:一种面向虚拟试穿的细粒度评估框架,将服装一致性分解为七个可解释维度——廓形、颜色、领口与袖型、主要装饰与结构、材质纹理、细节保真度、标志保留,每个维度均建模为离散属性预测任务。为训练该评估模型,采用两阶段学习范式:先在5万样本上进行大规模弱监督,再基于多模型投票获取1万高质量标注进行精调。同时使用加权交叉熵损失缓解各维度间的严重标签不平衡问题。该评估模型不仅可用于评价,还可集成至Qwen-Image-Edit的强化学习优化中,通过自适应聚合细粒度奖励,动态强调训练中表现欠佳的维度。实验表明,该方法(8B参数)在平衡准确率、SROCC和PLCC上达到当前最优,优于Gemini-3.1、Qwen3.7-plus、GPT-5.5等强大多模态模型,同时有效作为奖励信号指导生成优化。
原文摘要 · Abstract (English)
Virtual try-on (VTON) requires not only realistic generation but also faithful preservation of garment characteristics. However, existing evaluation metrics such as PSNR, SSIM, KID and FID struggle to measure the consistency between the generated and reference garments, particularly in capturing the multi-dimensional characteristics of garment fidelity. To address this, we propose DAT: a Dimension-wise Assessment framework for virtual Try-on, which decomposes garment consistency into seven interpretable dimensions: silhouette, color, neckline and sleeve shape, major decoration and structure, material texture, fine-detail fidelity, and logo preservation, each formulated as a discrete attribute-level prediction task. To train this specialized assessment model, we adopt a two-stage learning paradigm comprising large-scale weak supervision on 50K samples, followed by refinement on 10K higher-quality annotations obtained via multi-model voting. Furthermore, we employ weighted cross-entropy loss to mitigate the severe label imbalance inherent across evaluation dimensions. Beyond its role as an evaluation framework, the assessment model can be integrated into reinforcement learning optimization of Qwen-Image-Edit for VTON, where dimension-wise rewards are adaptively aggregated to emphasize under-optimized aspects during training. Experimental results show that our method (8B parameters) achieves state-of-the-art performance in terms of balanced accuracy, SROCC, and PLCC, outperforming strong proprietary models such as Gemini-3.1, Qwen3.7-plus, and GPT-5.5, while also serving as an effective optimization signal for reward-guided VTON generation
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。