arXiv:2507.18633cs.CV2025-07

从生成图像中识别被提及的艺术家,助力AI画作版权监管

Identifying Prompted Artist Names from Generated Images

  • 构建110位艺术家的195万张图像数据集,评估多种识别方法
  • 监督学习在已见艺术家上表现最好,但多艺术家提示仍难识别
  • 公开数据集与评测基准,推动生成内容责任治理研究

文本到图像模型常被用于明确指定艺术家风格(如“模仿格雷格·鲁特科夫斯基”)。本文提出一个受控艺术家识别基准:仅凭图像预测提示中提到的艺术家名称。数据集包含195万张图像,覆盖110位艺术家,涵盖四种泛化场景:未见艺术家、提示复杂度提升、多艺术家提示以及不同文本到图像模型。评估了特征相似性基线、对比风格描述符、数据归属方法、监督分类器和少样本原型网络。结果表明:监督与少样本模型在已见艺术家和复杂提示下表现优异,而风格描述符在风格明显时迁移性更强;多艺术家提示仍是最大挑战。该基准揭示了显著的改进空间,提供公开测试平台以推进文本到图像模型的负责任管控。数据集与评测工具已开源:https://graceduansu.github.io/IdentifyingPromptedArtists/

原文摘要 · Abstract (English)

A common and controversial use of text-to-image models is to generate pictures by explicitly naming artists, such as "in the style of Greg Rutkowski". We introduce a benchmark for prompted-artist recognition: predicting which artist names were invoked in the prompt from the image alone. The dataset contains 1.95M images covering 110 artists and spans four generalization settings: held-out artists, increasing prompt complexity, multiple-artist prompts, and different text-to-image models. We evaluate feature similarity baselines, contrastive style descriptors, data attribution methods, supervised classifiers, and few-shot prototypical networks. Generalization patterns vary: supervised and few-shot models excel on seen artists and complex prompts, whereas style descriptors transfer better when the artist's style is pronounced; multi-artist prompts remain the most challenging. Our benchmark reveals substantial headroom and provides a public testbed to advance the responsible moderation of text-to-image models. We release the dataset and benchmark to foster further research: https://graceduansu.github.io/IdentifyingPromptedArtists/

图像识别艺术风格版权监管数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。