构建10万张高清虚拟试衣数据集,提升评估准确性
OpenVTON-Bench: A Large-Scale High-Resolution Benchmark for Controllable Virtual Try-On Evaluation
- 用DINOv3聚类和Gemini生成描述,确保20类服装均衡分布
- 提出多模态评估协议,五维量化真实感、纹理等关键指标
- 新评测方法比传统指标更贴近真人评价,适合工业级应用
扩散模型显著提升了虚拟试衣系统的视觉质量,但可靠评估仍是瓶颈。传统指标难以捕捉细粒度纹理与语义一致性,现有数据集在规模与多样性上无法满足商业需求。我们提出OpenVTON-Bench,一个包含约10万对高分辨率图像(最高1536×1536)的大规模基准数据集。通过DINOv3分层聚类实现语义均衡采样,结合Gemini生成密集描述,确保20个细粒度服装类别分布均匀。为支持可靠评估,提出多模态评测协议,从背景一致性、身份保真度、纹理保真度、形状合理性及整体真实感五个可解释维度衡量性能。该协议融合VLM语义推理与基于SAM3分割和形态学腐蚀的新式多尺度表示度量,可分离边界对齐误差与内部纹理缺陷。实验显示其与人工判断高度一致(肯德尔τ达0.833,远超SSIM的0.611),建立了可靠的虚拟试衣评估基准。
原文摘要 · Abstract (English)
Recent advances in diffusion models have significantly elevated the visual fidelity of Virtual Try-On (VTON) systems, yet reliable evaluation remains a persistent bottleneck. Traditional metrics struggle to quantify fine-grained texture details and semantic consistency, while existing datasets fail to meet commercial standards in scale and diversity. We present OpenVTON-Bench, a large-scale benchmark comprising approximately 100K high-resolution image pairs (up to $1536 \times 1536$). The dataset is constructed using DINOv3-based hierarchical clustering for semantically balanced sampling and Gemini-powered dense captioning, ensuring a uniform distribution across 20 fine-grained garment categories. To support reliable evaluation, we propose a multi-modal protocol that measures VTON quality along five interpretable dimensions: background consistency, identity fidelity, texture fidelity, shape plausibility, and overall realism. The protocol integrates VLM-based semantic reasoning with a novel Multi-Scale Representation Metric based on SAM3 segmentation and morphological erosion, enabling the separation of boundary alignment errors from internal texture artifacts. Experimental results show strong agreement with human judgments (Kendall's $τ$ of 0.833 vs. 0.611 for SSIM), establishing a robust benchmark for VTON evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。