保护用户隐私的同时,让大模型回答更安全可靠。
PRIV-QA: Privacy-Preserving Question Answering for Cloud Large Language Models
- 构建多阶段隐私保护流程,提前拦截敏感信息泄露。
- 推出首个中英双语隐私问答数据集,含5.7万条交互数据。
- 适合关注云上大模型隐私安全的研究者与开发者。
大语言模型(LLMs)的快速发展正在重塑人机交互格局,其在各类用户服务应用中的集成日益普遍。然而,将用户数据传输至云端大模型存在数据泄露和敏感信息被非法访问的重大风险。本文提出一种隐私保护流水线,用于实际使用场景中用户与大模型交互时的隐私防护。我们构建了SensitiveQA——首个面向隐私保护的开放问答数据集,包含57,000条中英文交互数据,涵盖多样化的用户敏感信息。所提方案采用多阶段策略,在保障云端大模型响应质量的同时,主动保护用户信息。实验验证了该方法在隐私保护与交互质量之间取得良好平衡的有效性。代码与数据集已公开于https://github.com/ligw1998/PRIV-QA。
原文摘要 · Abstract (English)
The rapid development of large language models (LLMs) is redefining the landscape of human-computer interaction, and their integration into various user-service applications is becoming increasingly prevalent. However, transmitting user data to cloud-based LLMs presents significant risks of data breaches and unauthorized access to personal identification information. In this paper, we propose a privacy preservation pipeline for protecting privacy and sensitive information during interactions between users and LLMs in practical LLM usage scenarios. We construct SensitiveQA, the first privacy open-ended question-answering dataset. It comprises 57k interactions in Chinese and English, encompassing a diverse range of user-sensitive information within the conversations. Our proposed solution employs a multi-stage strategy aimed at preemptively securing user information while simultaneously preserving the response quality of cloud-based LLMs. Experimental validation underscores our method's efficacy in balancing privacy protection with maintaining robust interaction quality. The code and dataset are available at https://github.com/ligw1998/PRIV-QA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。