构建面向真实场景的虚拟试穿评估体系,全面检验模型表现。
VTBench: Comprehensive Benchmark Suite Towards Real-World Virtual Try-on Models
- 分层拆解试穿任务,覆盖图像质量、纹理保持等五维评估
- 引入真人偏好标注,确保评测结果贴近人类视觉感知
- 涵盖复杂户外场景,揭示模型在真实环境中的短板
尽管虚拟试穿技术取得显著进展,但在真实场景下的评估仍具挑战。现有评测存在三大不足:(1)指标难以反映人类感知,尤其在无配对场景下;(2)多数测试集局限于室内场景,缺乏真实复杂性;(3)缺乏能引导未来发展的系统化评估体系。为此,我们提出VTBench,一个分层的综合评测套件,将虚拟图像试穿任务分解为五个可解耦的关键维度:整体图像质量、纹理保持、复杂背景一致性、跨类别尺寸适应性与手部遮挡处理能力。每个维度配备定制化测试集与评估标准。该套件具备三大优势:(1)多维度评估框架:通过细粒度指标定位模型在多样化挑战场景中的优劣;(2)人眼对齐:所有测试集均提供真人偏好标注,确保评测与感知质量一致;(3)深度洞察:不仅涵盖常规室内场景,还分析模型在不同维度的表现差异,并探究室内与真实场景间的性能差距。为推动虚拟试穿向真实场景演进,VTBench将开源,包含全部测试集、评估协议、生成结果及人工标注。
原文摘要 · Abstract (English)
While virtual try-on has achieved significant progress, evaluating these models towards real-world scenarios remains a challenge. A comprehensive benchmark is essential for three key reasons:(1) Current metrics inadequately reflect human perception, particularly in unpaired try-on settings;(2)Most existing test sets are limited to indoor scenarios, lacking complexity for real-world evaluation; and (3) An ideal system should guide future advancements in virtual try-on generation. To address these needs, we introduce VTBench, a hierarchical benchmark suite that systematically decomposes virtual image try-on into hierarchical, disentangled dimensions, each equipped with tailored test sets and evaluation criteria. VTBench exhibits three key advantages:1) Multi-Dimensional Evaluation Framework: The benchmark encompasses five critical dimensions for virtual try-on generation (e.g., overall image quality, texture preservation, complex background consistency, cross-category size adaptability, and hand-occlusion handling). Granular evaluation metrics of corresponding test sets pinpoint model capabilities and limitations across diverse, challenging scenarios.2) Human Alignment: Human preference annotations are provided for each test set, ensuring the benchmark's alignment with perceptual quality across all evaluation dimensions. (3) Valuable Insights: Beyond standard indoor settings, we analyze model performance variations across dimensions and investigate the disparity between indoor and real-world try-on scenarios. To foster the field of virtual try-on towards challenging real-world scenario, VTBench will be open-sourced, including all test sets, evaluation protocols, generated results, and human annotations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。