用概率估计提升大模型查询代理的隐私保护能力
Beyond Direct Identifiers: Probabilistic Privacy Risk Estimation for Privacy-Conscious LLM Query Delegation

- 引入概率化k匿名性评估,识别隐性身份线索
- 在新数据集PUPA-SD上验证,优化后隐私-效果平衡更优
- 适合关注隐私安全的大模型应用开发者
现有隐私保护方法多聚焦于显式个人信息(PII),但用户仍可能通过非PII的自我披露信息被识别。本文研究一种概率化的隐私感知委托(PCD)机制,引入基于LLM的k匿名性概率估计作为辅助目标。为支持该研究,构建了包含自然对话中自我披露内容的PUPA-SD数据集。初步实验表明,在该数据集上优化PAPILLON模型能提升多种本地模型在未见对话上的表现,并为Llama-3.2-3B实现最佳隐私-效用平衡;而小型模型难以同时兼顾质量与隐私。研究提出将k匿名性作为有效辅助指标用于改善隐私保护。
原文摘要 · Abstract (English)
Recent work on protecting privacy during user-LLM interactions often focuses on direct, explicit identifiers: the personally-identifiable information (PII) captured by standard detectors. One such approach is Privacy-Conscious Delegation (PCD), where a local LLM acts as an intermediary. However, privacy risk does not stem solely from explicit identifiers but also PII-free self-disclosures, leaving users identifiable through combinations of quasi-identifying traits. We investigate a probabilistic variant of PCD, where we augment its objectives with an LLM-driven probabilistic estimation of k-anonymity. To facilitate this, we first create the PUPA-SD dataset, which contains naturalistic user queries with self-disclosure. Our preliminary results indicate that optimizing PAPILLON on PUPA-SD improves quality on unseen conversations across a variety of local models and produces the best privacy-utility balance for Llama-3.2-3B, while smaller models struggle to jointly optimize quality and privacy. We propose k-anonymity as a useful auxiliary metric for tackling PCD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。