构建超77万张虚拟试穿图像数据集,助力识别AI生成内容真伪。
VTONGuard: Automatic Detection and Authentication of AI-Generated Virtual Try-On Content
- 基于真实与合成图像构建大规模检测基准数据集。
- 多任务框架融合分割信息,显著提升边界特征识别能力。
- 适合研究AI内容安全、数字时尚及生成技术可信性的学者。
随着生成式AI的快速发展,虚拟试穿(VTON)系统在电商和数字娱乐中日益普及。然而,AI生成试穿内容的真实感不断提升,引发对真实性与负责任使用的担忧。为此,我们提出了VTONGuard,一个包含超过77.5万张真实与合成试穿图像的大规模基准数据集。该数据集涵盖多种真实场景条件,包括姿态、背景和服装风格的变化,并提供真实与篡改样本。基于此基准,我们在统一训练与测试协议下对多种检测范式进行了系统评估。结果揭示了各方法的优劣,凸显了跨范式泛化的持续挑战。为进一步提升检测性能,我们设计了一种多任务框架,通过引入辅助分割增强边界感知特征学习,在VTONGuard上取得最佳综合表现。我们期望该基准能实现公平比较,推动更鲁棒检测模型的发展,促进VTON技术在实际中的安全可靠部署。
原文摘要 · Abstract (English)
With the rapid advancement of generative AI, virtual try-on (VTON) systems are becoming increasingly common in e-commerce and digital entertainment. However, the growing realism of AI-generated try-on content raises pressing concerns about authenticity and responsible use. To address this, we present VTONGuard, a large-scale benchmark dataset containing over 775,000 real and synthetic try-on images. The dataset covers diverse real-world conditions, including variations in pose, background, and garment styles, and provides both authentic and manipulated examples. Based on this benchmark, we conduct a systematic evaluation of multiple detection paradigms under unified training and testing protocols. Our results reveal each method's strengths and weaknesses and highlight the persistent challenge of cross-paradigm generalization. To further advance detection, we design a multi-task framework that integrates auxiliary segmentation to enhance boundary-aware feature learning, achieving the best overall performance on VTONGuard. We expect this benchmark to enable fair comparisons, facilitate the development of more robust detection models, and promote the safe and responsible deployment of VTON technologies in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。