首个面向虚拟试穿的多维度质量评估数据集,解决生成图像失真问题。
VTONQA: A Multi-Dimensional Quality Assessment Dataset for Virtual Try-on
- 构建包含8132张图像的多维评估数据集,覆盖11种主流模型
- 通过24396个评分验证现有方法在服装贴合、身体匹配等维度的不足
- 适合研究虚拟试穿与图像质量评估的开发者和学者使用
随着电子商务和数字时尚的快速发展,基于图像的虚拟试穿(VTON)受到越来越多关注。然而,现有VTON模型常出现服装变形和人体不一致等伪影,亟需可靠的生成图像质量评估方法。为此,我们构建了首个专为VTON设计的多维度质量评估数据集VTONQA,包含11种代表性VTON模型生成的8,132张图像,以及在三个评价维度(服装贴合度、身体兼容性、整体质量)下的24,396个平均意见分数(MOS)。基于该数据集,我们对VTON模型及多种图像质量评估(IQA)指标进行了基准测试,揭示了现有方法的局限性,并凸显了本数据集的价值。我们相信,VTONQA及其基准将为感知对齐的评估提供坚实基础,推动质量评估方法与VTON模型的发展。数据集已公开:https://huggingface.co/datasets/weixiny0408/vtonqa。
原文摘要 · Abstract (English)
With the rapid development of e-commerce and digital fashion, image-based virtual try-on (VTON) has attracted increasing attention. However, existing VTON models often suffer from artifacts such as garment distortion and body inconsistency, highlighting the need for reliable quality evaluation of VTON-generated images. To this end, we construct \textbf{VTONQA}, the first multi-dimensional quality assessment dataset specifically designed for VTON, which contains 8,132 images generated by 11 representative VTON models, along with 24,396 mean opinion scores (MOSs) across three evaluation dimensions (\textit{i.e.}, clothing fit, body compatibility, and overall quality). Based on VTONQA, we benchmark both VTON models and a diverse set of image quality assessment (IQA) metrics, revealing the limitations of existing methods and highlighting the value of the proposed dataset. We believe that the VTONQA dataset and corresponding benchmarks will provide a solid foundation for perceptually aligned evaluation, benefiting both the development of quality assessment methods and the advancement of VTON models. The dataset we proposed in this paper is publicly available at: https://huggingface.co/datasets/weixiny0408/vtonqa.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。