arXiv:2503.13872cs.LGcs.CR2025-03被引 3

用重建攻击替代成员推断,更有效校准语言模型的隐私保护强度。

Empirical Calibration and Metric Differential Privacy in Language Models

  • 以角度距离扰动替代高斯噪声,提出基于VMF分布的定向隐私机制
  • 重建攻击比成员推断更适合作为隐私校准工具,实测验证其有效性
  • 发现不同隐私机制在隐私-效用权衡上各有优势,适用于不同场景

采用差分隐私训练的NLP模型通常使用DP-SGD框架,隐私保障常以隐私预算ε表示。但ε无内在意义,难以跨框架比较。图像领域已有研究通过成员推断攻击(MIA)实现噪声的实证校准,但该方法尚未应用于NLP。本文表明,MIA对隐私校准帮助甚微,而重建攻击更为有效。作为应用案例,我们提出一种基于冯·米塞斯-费舍尔(VMF)分布的新型方向性隐私机制,该机制扰动向量间角度距离而非添加各向同性高斯噪声,并应用于NLP架构。尽管形式化保证不可直接比较,但实证隐私校准显示,各机制在隐私-效用权衡中表现各异,具备不同适用优势。

原文摘要 · Abstract (English)

NLP models trained with differential privacy (DP) usually adopt the DP-SGD framework, and privacy guarantees are often reported in terms of the privacy budget $ε$. However, $ε$ does not have any intrinsic meaning, and it is generally not possible to compare across variants of the framework. Work in image processing has therefore explored how to empirically calibrate noise across frameworks using Membership Inference Attacks (MIAs). However, this kind of calibration has not been established for NLP. In this paper, we show that MIAs offer little help in calibrating privacy, whereas reconstruction attacks are more useful. As a use case, we define a novel kind of directional privacy based on the von Mises-Fisher (VMF) distribution, a metric DP mechanism that perturbs angular distance rather than adding (isotropic) Gaussian noise, and apply this to NLP architectures. We show that, even though formal guarantees are incomparable, empirical privacy calibration reveals that each mechanism has different areas of strength with respect to utility-privacy trade-offs.

差分隐私语言模型隐私校准重建攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。