用语义一致的假名保护对话隐私,不降质还防泄露
Semantically-Aware LLM Agent to Enhance Privacy in Conversational AI Services
- 动态替换敏感信息为语义匹配的假名,保持对话连贯性
- 相比基线方法,语义误判率降低5倍,隐私保护更彻底
- 适合需要高隐私保障的医疗、金融类对话场景
随着对话式AI系统广泛应用,用户在与大语言模型(LLMs)交互时分享敏感个人信息引发隐私泄露担忧。对话中可能包含个人身份信息(PII),一旦暴露可能导致安全风险或身份盗用。为此,我们提出局部伪名化语义保真实体检测框架(LOPSIDED),一种语义感知的隐私代理,用于保护远程LLM使用中的敏感数据。该方法动态将用户提示中的敏感PII实体替换为语义一致的伪名,保持对话上下文完整性;模型生成回复后,伪名自动还原,确保用户获得准确且隐私安全的输出。我们在来自ShareGPT的真实对话基础上进行增强与标注,评估命名实体是否与模型响应语境相关。结果表明,相较于基线技术,LOPSIDED将语义效用错误减少5倍,同时显著提升隐私保护能力。
原文摘要 · Abstract (English)
With the increasing use of conversational AI systems, there is growing concern over privacy leaks, especially when users share sensitive personal data in interactions with Large Language Models (LLMs). Conversations shared with these models may contain Personally Identifiable Information (PII), which, if exposed, could lead to security breaches or identity theft. To address this challenge, we present the Local Optimizations for Pseudonymization with Semantic Integrity Directed Entity Detection (LOPSIDED) framework, a semantically-aware privacy agent designed to safeguard sensitive PII data when using remote LLMs. Unlike prior work that often degrade response quality, our approach dynamically replaces sensitive PII entities in user prompts with semantically consistent pseudonyms, preserving the contextual integrity of conversations. Once the model generates its response, the pseudonyms are automatically depseudonymized, ensuring the user receives an accurate, privacy-preserving output. We evaluate our approach using real-world conversations sourced from ShareGPT, which we further augment and annotate to assess whether named entities are contextually relevant to the model's response. Our results show that LOPSIDED reduces semantic utility errors by a factor of 5 compared to baseline techniques, all while enhancing privacy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。