arXiv:2605.22373cs.LGcs.CL2026-05

针对安全分类器的边界攻击可高效识别敏感训练数据。

Boundary-targeted Membership Inference Attacks on Safety Classifiers

论文配图:Boundary-targeted Membership Inference Attacks on Safety Classifiers
图 1 · 摘自论文原文
  • 聚焦模型低置信度样本,定位边界区域增强成员推断信号。
  • 在5%误报率下恢复19%的用户求助对话,提升3.5倍于现有方法。
  • 揭示内容过滤无效,噪声策略可有效缓解攻击风险。

安全分类器是生成式AI系统中的关键防护机制,用于过滤有害内容或识别有心理风险的用户。然而,这些模型基于包含自残与心理健康讨论等敏感数据进行训练,引发严重隐私担忧。成员推断攻击(MIAs)使攻击者可推测训练数据中是否包含特定样本。本文假设:模型最不确定的样本对成员推断最具信息量,反映局部泛化失败,模型依赖记忆解决训练集歧义。为此,我们提出一种边界目标选择策略,识别能放大样本成员身份信号的低置信度样本。实验表明,在一个用于检测可能需要情感支持用户的分类器上,攻击者可在5%假阳性率下恢复19%的被标记为用户困扰的对话,比当前最优MIA方法高出3.5倍。最后,我们分析了边界样本特征,发现基于内容的过滤无法提供保护,而现有噪声策略能有效降低此类样本的脆弱性。

原文摘要 · Abstract (English)

Safety classifiers are essential safeguards within generative AI systems, filtering harmful content or identifying at-risk users when interacting with large language models. Despite their necessity, these models are trained on sensitive datasets including discussions of self-harm and mental health, raising important, yet poorly understood, privacy concerns. Membership inference attacks (MIAs) allow adversaries to infer membership of examples used to train models. In this work, we hypothesize that identifying the examples on which the classifier is least confident are informative for an adversary to infer membership. This reflects a localized failure of generalization, where the model relies on memorization to resolve ambiguity in the training set. To investigate this, we introduce a new boundary-targeted selection strategy that identifies low confidence examples that amplify the signal of an examples membership within a training set. Our experimental results show that an adversary can recover 19% of the conversations a safety classifier flagged as indicating user distress, at a 5% false-positive rate, on a classifier fine-tuned for detecting a user who may require emotional support. This is $3.5$ times more than attacking using state-of-the-art MIA methods alone. Finally, we characterize the boundary laying examples and show that content-based filtering is ineffective for protection, and existing noise strategies can effectively mitigate susceptibility of these examples.

隐私安全成员推断大模型防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。