提出新框架,精准识别RAG系统中的隐私泄露风险
Privacy Policy Enforcement Guardrails for Data-Sensitive Retrieval-Augmented Generation

- 用双单类密度估计融合文本嵌入,检测隐私违规
- 边界安全测试下AUROC达0.93+,误报率降低44%-55%
- 适合需低延迟高可靠性的数据敏感场景
标准的个人身份信息(PII)过滤器常遗漏RAG系统中的上下文数据泄露问题,例如未受监管的属性簇组合仍可能识别个体。本文提出隐私政策执行(PPE)框架,采用双单类密度估计器结合融合文本嵌入,并设置校准后的拒绝区域以应对分布外输入。通过在医疗、金融、法律领域构建轴向分层的多大模型合成数据流水线,发现传统高斯混合基线在边界安全压力测试中表现不佳,因其过度关注语言风格而非内容实质。所提T3+OCSVM检测器在安全与边界安全数据上训练,实现边界状态下AUROC超过0.93,同时将误报率降低44-55个百分点,并保持毫秒级延迟。相比监督式MLP分类器或140亿参数大模型判断器,本框架在操作适用性上更优——前者存在高拒答率,后者则面临延迟与校准问题。该方法为任何基于合成数据训练的分类器提供了稳健的压力测试标准。
原文摘要 · Abstract (English)
Standard PII filters often miss contextual data leakage in RAG systems, such as non-regulated attribute clusters that collectively identify individuals. We introduce a Privacy Policy Enforcement (PPE) framework using dual one-class density estimators with fused text embeddings and a calibrated abstain region for out-of-distribution inputs. Using an axis-stratified, multi-LLM synthetic data pipeline across medicine, finance, and law, we found that traditional Gaussian Mixture baselines fail on borderline-safe stress tests by focusing on linguistic register rather than content. Our proposed T3+OCSVM detector, trained on safe and borderline-safe data, achieves a borderline AUROC of 0.93+ while reducing false positives by 44-55 percentage points and maintaining millisecond latency. Compared to supervised MLP classifiers or 14B-parameter LLM judges, our framework offers superior operational suitability, as the former suffers from high abstention rates and the latter from latency and calibration issues. This methodology provides a robust stress-testing standard for any synthetic-data-trained classifier.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。