arXiv:2410.17127cs.CRcs.CL2024-10NAACL被引 63

用本地+云端模型协作,保护隐私同时保持高回答质量

PAPILLON: Privacy Preservation from Internet-based and Local Language Model Ensembles

  • 设计多阶段模型管道,通过优化提示实现隐私优先的混合推理
  • 85.5%查询保持高质量响应,仅7.5%存在隐私泄露风险
  • 适用于关注隐私安全的智能助手、企业级应用开发

用户向专有大模型提供商披露敏感信息,引发重大隐私担忧。尽管开源本地模型可缓解部分顾虑,但本地部署模型的能力通常弱于专有前沿模型。为在保障用户隐私的同时保留最佳性能,我们提出隐私感知委托(Privacy-Conscious Delegation)这一新任务,即串联基于API的模型与本地模型。我们利用近期公开的用户-大模型交互数据集构建自然基准PUPA,其中包含个人身份信息(PII)。为探索可行方案,我们设计PAPILLON,一种多阶段大模型流水线,通过提示优化应对该任务的简化版本。最优流水线在85.5%的用户查询中维持高响应质量,同时将隐私泄露控制在7.5%。未来工作仍需提升生成质量以逼近专有模型水平。数据与代码已公开于https://github.com/siyan-sylvia-li/PAPILLON。

原文摘要 · Abstract (English)

Users can divulge sensitive information to proprietary LLM providers, raising significant privacy concerns. While open-source models, hosted locally on the user's machine, alleviate some concerns, models that users can host locally are often less capable than proprietary frontier models. Toward preserving user privacy while retaining the best quality, we propose Privacy-Conscious Delegation, a novel task for chaining API-based and local models. We utilize recent public collections of user-LLM interactions to construct a natural benchmark called PUPA, which contains personally identifiable information (PII). To study potential approaches, we devise PAPILLON, a multi-stage LLM pipeline that uses prompt optimization to address a simpler version of our task. Our best pipeline maintains high response quality for 85.5% of user queries while restricting privacy leakage to only 7.5%. We still leave a large margin to the generation quality of proprietary LLMs for future work. Our data and code is available at https://github.com/siyan-sylvia-li/PAPILLON.

隐私保护模型协同本地部署提示优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。