提出无需训练的隐私保护机制ABack,防范大模型企业数据泄露。
Adaptive Backtracking for Privacy Protection in Large Language Models
- 用隐状态模型定位泄露意图,安全重写输出内容。
- 在医疗金融场景下,隐私保护效果比基线提升15%。
- 构建新数据集PriGenQA,支持对抗性攻击评估。
隐私保护已成为人工智能时代的关键议题。现有研究多关注用户隐私,忽视了检索增强生成范式加剧的企业数据泄露风险。为此,本文提出面向企业的隐私保护目标,解决两大挑战:现有数据清洗方法严重损害模型性能,且缺乏公开评估数据集。为此,我们提出ABack——一种无需训练的机制,通过隐状态模型识别泄露意图并安全重写输出;构建了面向医疗与金融场景的私密性评估基准PriGenQA;并设计基于群体相对策略优化的自适应攻击者,实现更严格的评估。实验表明,在该强对抗环境下,ABack相较强基线隐私效用得分提升最高达15%,避免了以往方法的性能损失。
原文摘要 · Abstract (English)
The preservation of privacy has emerged as a critical topic in the era of artificial intelligence. However, current work focuses on user-oriented privacy, overlooking severe enterprise data leakage risks exacerbated by the Retrieval-Augmented Generation paradigm. To address this gap, our paper introduces a novel objective: enterprise-oriented privacy concerns. Achieving this objective requires overcoming two fundamental challenges: existing methods such as data sanitization severely degrade model performance, and the field lacks public datasets for evaluation. We address these challenges with several solutions. (1) To prevent performance degradation, we propose ABack, a training-free mechanism that leverages a Hidden State Model to pinpoint the origin of a leakage intention and rewrite the output safely. (2) To solve the lack of datasets, we construct PriGenQA, a new benchmark for enterprise privacy scenarios in healthcare and finance. To ensure a rigorous evaluation, we move beyond simple static attacks by developing a powerful adaptive attacker with Group Relative Policy Optimization. Experiments show that against this superior adversary, ABack improves the overall privacy utility score by up to 15\% over strong baselines, avoiding the performance trade-offs of prior methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。