arXiv:2510.16122cs.CRcs.CL2025-10

生成式文本分类器易遭成员推理攻击,存在隐私泄露风险。

The Hidden Cost of Modeling P(X): Vulnerability to Membership Inference Attacks in Generative Text Classifiers

  • 通过建模联合概率P(X,Y)的生成式分类器更易泄露训练数据成员身份。
  • 在9个基准数据集上,生成式模型的成员推理成功率显著高于判别式模型。
  • 适合关注生成模型隐私安全的研究者与应用开发者参考。

成员推理攻击(MIA)能迫使攻击者判断特定样本是否曾被用于模型训练,构成严重隐私威胁。尽管已有大量关于MIA的研究,但对生成式与判别式分类器的系统性比较仍不足。本文首先从理论上解释为何生成式分类器更易受MIA影响,随后在九个基准数据集上,针对不同训练规模的判别式、生成式及伪生成式文本分类器进行了全面实证评估。采用多种MIA策略,结果一致表明:显式建模联合分布P(X,Y)的完整生成式分类器最易遭受成员信息泄露;此外,生成式分类器中常见的经典推理方式进一步放大了隐私风险。研究揭示了分类器设计中的根本性效用-隐私权衡,警示在隐私敏感场景部署生成式模型需谨慎。结果为未来构建兼具性能与隐私保护能力的生成式分类器提供了方向。

原文摘要 · Abstract (English)

Membership Inference Attacks (MIAs) pose a critical privacy threat by enabling adversaries to determine whether a specific sample was included in a model's training dataset. Despite extensive research on MIAs, systematic comparisons between generative and discriminative classifiers remain limited. This work addresses this gap by first providing theoretical motivation for why generative classifiers exhibit heightened susceptibility to MIAs, then validating these insights through comprehensive empirical evaluation. Our study encompasses discriminative, generative, and pseudo-generative text classifiers across varying training data volumes, evaluated on nine benchmark datasets. Employing a diverse array of MIA strategies, we consistently demonstrate that fully generative classifiers which explicitly model the joint likelihood $P(X,Y)$ are most vulnerable to membership leakage. Furthermore, we observe that the canonical inference approach commonly used in generative classifiers significantly amplifies this privacy risk. These findings reveal a fundamental utility-privacy trade-off inherent in classifier design, underscoring the critical need for caution when deploying generative classifiers in privacy-sensitive applications. Our results motivate future research directions in developing privacy-preserving generative classifiers that can maintain utility while mitigating membership inference vulnerabilities.

隐私安全生成模型成员推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。