arXiv:2506.05126cs.CRcs.LG2025-06中稿 · the 8th Deep Learn…被引 9

针对序列模型的隐私泄露,提出基于序列相关性的会员推理攻击方法。

Membership Inference Attacks on Sequence Models

  • 利用序列生成中的内在相关性改进会员推理攻击
  • 在不增加计算成本下显著提升记忆审计效果
  • 适合关注大模型隐私安全的研究者与开发者

序列模型(如大型语言模型和自回归图像生成器)倾向于记忆并意外泄露敏感信息。尽管这一现象具有重大法律意义,但现有审计工具因假设不匹配而效果有限。我们提出,有效评估序列模型的隐私泄露需利用序列生成中的固有相关性。为此,我们改进了一种先进的会员推理攻击,显式建模序列内相关性,证明了该攻击可自然扩展以适应序列模型结构。案例研究表明,我们的改进在不增加计算开销的前提下,持续提升了记忆审计的有效性。本工作为大规模序列模型的可靠记忆审计提供了重要基础。

原文摘要 · Abstract (English)

Sequence models, such as Large Language Models (LLMs) and autoregressive image generators, have a tendency to memorize and inadvertently leak sensitive information. While this tendency has critical legal implications, existing tools are insufficient to audit the resulting risks. We hypothesize that those tools' shortcomings are due to mismatched assumptions. Thus, we argue that effectively measuring privacy leakage in sequence models requires leveraging the correlations inherent in sequential generation. To illustrate this, we adapt a state-of-the-art membership inference attack to explicitly model within-sequence correlations, thereby demonstrating how a strong existing attack can be naturally extended to suit the structure of sequence models. Through a case study, we show that our adaptations consistently improve the effectiveness of memorization audits without introducing additional computational costs. Our work hence serves as an important stepping stone toward reliable memorization audits for large sequence models.

隐私安全序列模型会员推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。