构建首个真实场景的合成图像溯源数据集,助力模型识别来源并抵抗后处理干扰。
WILD: a new in-the-Wild Image Linkage Dataset for synthetic image attribution
- 基于10个商用生成器构建封闭集,另增10个开放生成器模拟真实场景。
- 每类生成器各1000张图,共2万张,半数经多种后处理操作。
- 支持闭集、开集溯源及抗后处理、对抗攻击等多任务评估,适合安全与可信生成研究者。
合成图像来源溯源是一项开放挑战,随着每年大量图像生成工具发布,其复杂性与可用生成技术数量激增,且高质量、多样化的开源数据集稀缺,导致训练和评测溯源模型极为困难。WILD 是一个面向真实环境(in-the-Wild)的图像链接数据集,旨在为合成图像溯源模型提供强大训练与评测工具。该数据集由10个主流商业生成器构成封闭集作为模型训练基础,另加入10个额外生成器构成开放集,模拟真实世界中的未知来源场景。每个生成器提供1000张图像,封闭集与开放集各含10,000张图像。其中一半图像经过多种后处理操作。WILD 支持多种任务的评测,包括闭集与开集识别、验证以及对后处理和对抗攻击的鲁棒性评估。基于 WILD 训练的模型可受益于其高挑战性的现实场景设置。此外,论文还评估了七种基线方法在闭集与开集溯源任务上的表现,并进行了后处理鲁棒性测试。
原文摘要 · Abstract (English)
Synthetic image source attribution is an open challenge, with an increasing number of image generators being released yearly. The complexity and the sheer number of available generative techniques, as well as the scarcity of high-quality open source datasets of diverse nature for this task, make training and benchmarking synthetic image source attribution models very challenging. WILD is a new in-the-Wild Image Linkage Dataset designed to provide a powerful training and benchmarking tool for synthetic image attribution models. The dataset is built out of a closed set of 10 popular commercial generators, which constitutes the training base of attribution models, and an open set of 10 additional generators, simulating a real-world in-the-wild scenario. Each generator is represented by 1,000 images, for a total of 10,000 images in the closed set and 10,000 images in the open set. Half of the images are post-processed with a wide range of operators. WILD allows benchmarking attribution models in a wide range of tasks, including closed and open set identification and verification, and robust attribution with respect to post-processing and adversarial attacks. Models trained on WILD are expected to benefit from the challenging scenario represented by the dataset itself. Moreover, an assessment of seven baseline methodologies on closed and open set attribution is presented, including robustness tests with respect to post-processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。