构建首个多参考虚拟试穿评估基准,评测通用编辑模型表现
VTEdit-Bench: A Comprehensive Benchmark for Multi-Reference Image Editing Models in Virtual Try-On
- 设计包含24,220张图像对的跨场景评测集,覆盖5类渐进复杂度任务
- 提出VTEdit-QA评估框架,从一致性、衣物匹配度和图像质量三方面量化评价
- 发现顶尖通用模型在复杂场景下仍难处理多服饰参考条件
随着虚拟试穿(VTON)技术的发展,现实应用场景日益复杂,现有专用模型已难以应对。与此同时,通用多参考图像编辑模型快速发展,在视觉编辑中表现出强大泛化能力,为更灵活的VTON系统提供了新路径。然而,由于缺乏系统的评估基准,这些通用编辑器在VTON中的优势与局限仍不明确。为此,本文提出VTEdit-Bench,一个综合性基准,用于评估通用多参考图像编辑模型在多种真实VTON场景下的表现。该基准包含24,220个测试图像对,涵盖五类代表性任务,复杂度逐级递增,支持对鲁棒性和泛化能力的系统分析。我们进一步提出VTEdit-QA——一种基于视觉语言模型的参考感知评估器,从模型一致性、衣物一致性及整体图像质量三个维度评估性能。通过该框架,我们系统评估了八种通用编辑模型,并与七种专用VTON模型进行对比。结果表明,顶级通用编辑器在传统任务上表现竞争力,且在更复杂场景中泛化更稳定,但在多服饰条件等复杂参考配置下仍面临挑战。
原文摘要 · Abstract (English)
As virtual try-on (VTON) continues to advance, a growing number of real-world scenarios have emerged, pushing beyond the ability of the existing specialized VTON models. Meanwhile, universal multi-reference image editing models have progressed rapidly and exhibit strong generalization in visual editing, suggesting a promising route toward more flexible VTON systems. However, despite their strong capabilities, the strengths and limitations of universal editors for VTON remain insufficiently explored due to the lack of systematic evaluation benchmarks. To address this gap, we introduce VTEdit-Bench, a comprehensive benchmark designed to evaluate universal multi-reference image editing models across various realistic VTON scenarios. VTEdit-Bench contains 24,220 test image pairs spanning five representative VTON tasks with progressively increasing complexity, enabling systematic analysis of robustness and generalization. We further propose VTEdit-QA, a reference-aware VLM-based evaluator that assesses VTON performance from three key aspects: model consistency, cloth consistency, and overall image quality. Through this framework, we systematically evaluate eight universal editing models and compare them with seven specialized VTON models. Results show that top universal editors are competitive on conventional tasks and generalize more stably to harder scenarios, but remain challenged by complex reference configurations, particularly multi-cloth conditioning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。