arXiv:2605.04116cs.CRcs.LG2026-05

检索增强的上下文学习易受成员推理攻击,可暴露数据隐私。

Membership Inference Attacks for Retrieval Based In-Context Learning for Document Question Answering

论文配图:Membership Inference Attacks for Retrieval Based In-Context Learning for Document Question Answering
图 1 · 摘自论文原文
  • 利用查询前缀设计黑盒攻击,无需访问模型内部参数。
  • 新方法通过加权平均计算成员统计量,无需参考模型且更鲁棒。
  • 对改写文本仍有效,适合关注大模型隐私泄露的研究者。

我们证明,当远程托管的应用程序在上下文学习中引入检索功能以选择上下文示例时,即使服务提供方与用户分离,也可能面临成员推理攻击。本文提出两种黑盒攻击方法,利用查询文本前缀区分成员与非成员输入。第一种攻击借助参考模型估计不可用的损失指标;第二种攻击则消除参考模型,采用新颖的加权平均方案直接计算成员统计量。全面实验证明,在对手拥有查询文本改写版本的严格场景下,本方法对改写具有更强鲁棒性,在小前缀数量下多数情况下优于三种已有攻击。此外,我们将一种现有集成提示防御方法适配至本场景,证实其能显著缓解第二种攻击引发的隐私泄露。

原文摘要 · Abstract (English)

We show that remotely hosted applications employing in-context learning when augmented with a retrieval function to select in-context examples can be vulnerable to membership-inference attacks even when the service provider and users are separate parties. We propose two black-box membership inference attacks that exploit query text prefixes to distinguish member from non-member inputs. The first attack uses a reference model to estimate an otherwise unavailable loss metric. The second attack improves upon it by eliminating the reference model and instead computing a membership statistic through a simple but novel weighted-averaging scheme. Our comprehensive empirical evaluations consider a stricter case in which the adversary has a paraphrased version of the text in the queries and show that our attacks can exhibit stronger resilience to paraphrasing and outperform three prior attacks in many cases with small number of prefixes. We also adapt an existing ensemble prompting defense to our setting, demonstrating that it substantially mitigates the privacy leakage caused by our second attack.

隐私安全成员推理检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。