arXiv:2510.05582cs.LG2025-10被引 2

提出更精准的隐私泄露检测方法,可定位模型记住的具体词元

(Token-Level) InfoRMIA: Stronger Membership Inference and Memorization Assessment for LLMs

  • 基于信息论构建新型会员推断框架,比现有方法更有效
  • 首次实现从词元层面定位模型记忆内容,识别精度显著提升
  • 适合关注大模型隐私安全与数据遗忘的开发者与研究者

机器学习模型会不可避免地记忆训练数据中的敏感信息,尤其是当前大型语言模型(LLMs)几乎使用全部可用数据进行训练,加剧了信息泄露风险。因此,在发布前量化隐私风险至关重要。目前标准方法是会员推断攻击,其中最先进的方法为鲁棒会员推断攻击(RMIA)。本文提出InfoRMIA,一种基于信息论的会员推断新范式,其在多个基准测试中持续优于RMIA,同时具备更高计算效率。此外,我们指出将序列级会员推断视为金标准存在局限性,提出新的分析视角:以词元级信号为基础。结果显示,简单的词元级InfoRMIA能精确定位生成输出中被记忆的特定词元,实现了从序列级到词元级的泄漏定位,且在序列级推断性能上更强。该新范式重新定义了大模型隐私评估方式,有助于实现更精准的遗忘机制。

原文摘要 · Abstract (English)

Machine learning models are known to leak sensitive information, as they inevitably memorize (parts of) their training data. More alarmingly, large language models (LLMs) are now trained on nearly all available data, which amplifies the magnitude of information leakage and raises serious privacy risks. Hence, it is more crucial than ever to quantify privacy risk before the release of LLMs. The standard method to quantify privacy is via membership inference attacks, where the state-of-the-art approach is the Robust Membership Inference Attack (RMIA). In this paper, we present InfoRMIA, a principled information-theoretic formulation of membership inference. Our method consistently outperforms RMIA across benchmarks while also offering improved computational efficiency. In the second part of the paper, we identify the limitations of treating sequence-level membership inference as the gold standard for measuring leakage. We propose a new perspective for studying membership and memorization in LLMs: token-level signals and analyses. We show that a simple token-based InfoRMIA can pinpoint which tokens are memorized within generated outputs, thereby localizing leakage from the sequence level down to individual tokens, while achieving stronger sequence-level inference power on LLMs. This new scope rethinks privacy in LLMs and can lead to more targeted mitigation, such as exact unlearning.

隐私安全模型记忆会员推断大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。