arXiv:2607.15084cs.CV2026-07中稿 · IEEE/IAPR IJCB 202…

研究人脸模型中训练数据成员的身份可识别性,发现训练样本越多,越难区分是否在训练集中。

Quantifying Training Membership Information in the Hyperspherical Embedding Geometry of Face Recognition Models

  • 通过超球面几何分析模型对训练/非训练身份的区分能力
  • 训练身份数量是影响区分度的最关键因素,且效果随数量增加而单调下降
  • 跨域数据会夸大成员信号,适合关注隐私泄露风险的研究者

人脸识别模型将每张人脸映射到单位超球面上,通过角度间隔损失使同一身份的嵌入聚类,不同身份相互分离。由于这些损失仅作用于训练身份,非成员身份可能形成具有不同几何特征的簇。本文通过180个IResNet模型在骨干网络大小、损失头、训练时长和训练身份数量的因子设计下,计算四类基于簇几何的统计量,并在九个基准上评估。结果表明:训练身份数量对成员与非成员可分性影响最大,骨干网络和损失头影响较小;在同域保留测试集上,随着训练身份增加,几何成员信号单调减弱。我们还分析了跨域(姿态、年龄、质量、族裔)非成员基准,发现其会放大表面成员信号。最后,通过学习分类器融合四个统计量,揭示了比单个统计量更多成员信息。

原文摘要 · Abstract (English)

Face recognition models represent each face as an embedding vector on the unit hypersphere by clustering embeddings of the same identity while pushing different identities apart through angular-margin losses. Because these losses act only on training identities, non-member identities may form clusters with different geometric properties. In this paper, we quantify the magnitude of this difference and what training-time factors control it. We compute four statistics based on cluster geometry across 180 face recognition models in a factorial design over IResNet backbone size, loss head, training duration, and the number of training identities, and evaluate each configuration on nine benchmarks. Our results indicate that the number of training identities has the largest effect on member/non-member separability, while backbone and loss head contribute far less, and that, on a same-domain held-out reference, the geometric membership signal decreases monotonically as more identities are added to training. We provide an analysis of cross-domain (pose, age, quality, ethnicity) non-member benchmarks and report that these inflate the apparent membership signal. Finally, we fuse all four statistics with a learned classifier to reveal additional membership information beyond the best individual statistic.

人脸识别嵌入几何隐私泄露成员推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。