arXiv:2510.06525cs.LGcs.CR2025-10中稿 · NeurIPS

Text-to-image模型生成图像可被轻易识别,威胁排行榜安全。

Text-to-Image Models Leave Identifiable Signatures: Implications for Leaderboard Security

  • 用CLIP嵌入空间实时分类,无需提示控制即可识别模型
  • 15万张图、19个模型中识别准确率极高,部分提示下近乎完美
  • 揭示文本生成图像榜单易被排名操纵,需更强防护

生成式AI排行榜是评估模型能力的核心,但易受操控。其中关键威胁为排名操纵,攻击者需先对展示的输出进行模型去匿名化——这一问题此前已在大语言模型中被验证。本文表明,文本到图像排行榜的去匿名化问题更为严重,识别难度显著更低。基于280个提示和19种来自不同组织、架构与规模的模型生成的超过15万张图像,我们证明在CLIP嵌入空间中进行简单实时分类即可高精度识别生成模型,且无需提示控制或历史数据。进一步提出提示级可分性度量,识别出能实现近乎完美去匿名化的特定提示。结果表明,文本到图像排行榜的排名操纵比先前认知更易实现,亟需更强防御机制。

原文摘要 · Abstract (English)

Generative AI leaderboards are central to evaluating model capabilities, but remain vulnerable to manipulation. Among key adversarial objectives is rank manipulation, where an attacker must first deanonymize the models behind displayed outputs -- a threat previously demonstrated and explored for large language models (LLMs). We show that this problem can be even more severe for text-to-image leaderboards, where deanonymization is markedly easier. Using over 150,000 generated images from 280 prompts and 19 diverse models spanning multiple organizations, architectures, and sizes, we demonstrate that simple real-time classification in CLIP embedding space identifies the generating model with high accuracy, even without prompt control or historical data. We further introduce a prompt-level separability metric and identify prompts that enable near-perfect deanonymization. Our results indicate that rank manipulation in text-to-image leaderboards is easier than previously recognized, underscoring the need for stronger defenses.

生成模型模型指纹安全评测图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。