用互联网海量未标注数据提升生成图像溯源能力
Unknown Aware AI-Generated Content Attribution
- 基于CLIP特征与线性分类器,仅需少量目标模型数据建立基准
- 引入未标注网络数据后,在未知生成模型上准确率显著提升
- 适合需要应对新出现生成模型的检测系统开发者
随着高保真生成模型的快速发展,识别合成内容来源变得愈发重要,需从简单的真伪二分类转向具体模型溯源。本文研究如何区分目标生成模型(如OpenAI Dalle 3)输出与其他来源图像(包括真实图像及多种替代模型生成内容)的能力。利用CLIP特征与简单线性分类器,仅需少量目标模型标签数据和少数已知生成模型样本,即可建立有效基线。然而该方法在面对未知、新发布生成模型时泛化能力不足。为此,本文提出一种约束优化方法,利用互联网收集的未标注数据(可能包含真实图像、未知生成模型输出或目标模型样本),强制使这些野数据被分类为非目标模型,同时严格保持对标注数据的高分类性能。实验表明,引入野数据可显著提升对挑战性未知生成模型的溯源准确率,证明在开放世界中可有效利用未标注数据增强生成内容溯源能力。
原文摘要 · Abstract (English)
The rapid advancement of photorealistic generative models has made it increasingly important to attribute the origin of synthetic content, moving beyond binary real or fake detection toward identifying the specific model that produced a given image. We study the problem of distinguishing outputs from a target generative model (e.g., OpenAI Dalle 3) from other sources, including real images and images generated by a wide range of alternative models. Using CLIP features and a simple linear classifier, shown to be effective in prior work, we establish a strong baseline for target generator attribution using only limited labeled data from the target model and a small number of known generators. However, this baseline struggles to generalize to harder, unseen, and newly released generators. To address this limitation, we propose a constrained optimization approach that leverages unlabeled wild data, consisting of images collected from the Internet that may include real images, outputs from unknown generators, or even samples from the target model itself. The proposed method encourages wild samples to be classified as non target while explicitly constraining performance on labeled data to remain high. Experimental results show that incorporating wild data substantially improves attribution performance on challenging unseen generators, demonstrating that unlabeled data from the wild can be effectively exploited to enhance AI generated content attribution in open world settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。