arXiv:2606.20155cs.CVcs.CL2026-06

提出黑箱方法检测文生图模型是否记住名人身份,无需原始照片或训练数据。

NAMESAKES: Probing Identity Memorization in Text-to-Image Models

论文配图:NAMESAKES: Probing Identity Memorization in Text-to-Image Models
图 1 · 摘自论文原文
  • 设计全黑箱行为探测器,仅通过输入名称观察生成结果判断记忆情况。
  • 在超千名名人数据集上验证,能有效区分被记住与未被记住的姓名。
  • 适用于隐私评估,尤其适合研究模型记忆风险的研究者和开发者。

文生图(T2I)模型在提示特定人名时可能生成其逼真肖像,引发隐私担忧。然而,当前区分生成图像是否为模型记忆内容,通常需要真实参考图、训练数据访问权限或对模型内部的白盒访问,限制了实际应用。本文提出一种完全黑箱的行为探测方法,可在无需参考照片或训练数据先验知识的情况下,区分被记住与未被记住的人名。为评估该任务,我们构建了NAMESAKES数据集,包含超过一千个公众人物的姓名与人脸,涵盖广泛知名度,并引入经过扰动的较不知名姓名作为对照。在多个先进T2I模型上的实验表明,该探测器能显著预测身份记忆现象,并有效分离被记住与未被记住的姓名,同时揭示不同模型家族间的差异。

原文摘要 · Abstract (English)

Text-to-image (T2I) models generate realistic likenesses of some individuals when prompted with their names, raising privacy concerns. However, distinguishing whether a generated face is memorized or fabricated currently requires ground-truth photos, access to training data, or white-box access to model internals, limiting applicability. We introduce a fully black-box behavioral probe that distinguishes between memorized and unrecognized names, while requiring no reference photos or prior knowledge of training data. To benchmark this task, we present the NAMESAKES dataset of over one thousand names and faces of public figures spanning a wide range of fame levels, along with perturbed, less famous names. Experiments on state-of-the-art T2I models show that our probe substantially predicts identity memorization and separates memorized from unrecognized names, with further insights into differences across model families.

隐私安全模型记忆文生图黑箱探测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。