arXiv:2601.09647cs.CVcs.CR2026-01

破解文本生成图像排行榜的匿名性,发现模型可被精准识别

Identifying Models Behind Text-to-Image Leaderboards

  • 通过图像嵌入空间中的聚类特征,无需提示词控制即可识别模型
  • 使用22个模型、280个提示词生成15万张图像,识别准确率高
  • 揭示排行榜安全漏洞,适合关注AI安全与模型隐私的研究者

文本到图像(T2I)模型日益流行,占据了网络上大量AI生成图像。为公平比较模型质量,投票式排行榜成为标准,依赖匿名模型输出。本文表明,这种匿名性极易被突破。我们发现,每个T2I模型生成的图像在图像嵌入空间中形成独特聚类,可在无提示词控制或训练数据的情况下实现精准去匿名化。基于22个模型和280个提示词(共15万张图像)的中心点方法达到高识别准确率,并引入提示级可区分性度量。大规模分析显示,某些提示词可导致近乎完美的模型区分。研究揭示了T2I排行榜的根本安全缺陷,推动更强的匿名化防护措施。

原文摘要 · Abstract (English)

Text-to-image (T2I) models are increasingly popular, producing a large share of AI-generated images online. To compare model quality, voting-based leaderboards have become the standard, relying on anonymized model outputs for fairness. In this work, we show that such anonymity can be easily broken. We find that generations from each T2I model form distinctive clusters in the image embedding space, enabling accurate deanonymization without prompt control or training data. Using 22 models and 280 prompts (150K images), our centroid-based method achieves high accuracy and reveals systematic model-specific signatures. We further introduce a prompt-level distinguishability metric and conduct large-scale analyses showing how certain prompts can lead to near-perfect distinguishability. Our findings expose fundamental security flaws in T2I leaderboards and motivate stronger anonymization defenses.

图像生成模型识别安全漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。