估算文本生成图像模型模仿特定概念所需的最少图片数。
How Many Images Does It Take? Estimating Imitation Thresholds in Text-to-Image Models
- 提出高效方法,无需重新训练即可估算模仿阈值。
- 实验发现模仿阈值在200-700张图之间,因领域和模型而异。
- 为版权侵权主张提供实证依据,适合模型开发者参考。
文本到图像模型使用互联网上大规模的图像-文本对数据集进行训练,这些数据集常包含受版权保护或私密图像。在这些数据上训练模型可能导致生成与训练图像具有可识别相似性的图像,这种现象称为模仿。本文提出估算模型达到模仿某概念所需训练样本数量——模仿阈值的新问题,并提出一种无需从头训练即可高效估算的方法。我们在人脸和艺术风格两个领域,评估了四种在三个预训练数据集上训练的文本到图像模型。结果表明,模仿阈值范围为200至700张图像,具体取决于领域和模型。该阈值为版权侵权指控提供了实证基础,也为希望遵守版权和隐私法律的模型开发者提供了指导原则。
原文摘要 · Abstract (English)
Text-to-image models are trained using large datasets of image-text pairs collected from the internet. These datasets often include copyrighted and private images. Training models on such datasets enables them to generate images that might violate copyright laws and individual privacy. This phenomenon is termed imitation -- generation of images with content that has recognizable similarity to its training images. In this work we estimate the point at which a model was trained on enough instances of a concept to be able to imitate it -- the imitation threshold. We posit this question as a new problem and propose an efficient approach that estimates the imitation threshold without incurring the colossal cost of training these models from scratch. We experiment with two domains -- human faces and art styles, and evaluate four text-to-image models that were trained on three pretraining datasets. We estimate the imitation threshold of these models to be in the range of 200-700 images, depending on the domain and the model. The imitation threshold provides an empirical basis for copyright violation claims and acts as a guiding principle for text-to-image model developers that aim to comply with copyright and privacy laws. Website: https://how-many-van-goghs-does-it-take.github.io/. Code: https://github.com/vsahil/MIMETIC-2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。