arXiv:2512.24848cs.CLcs.AI2025-12被引 10

测试个性化AI泄露用户隐私的漏洞,发现26.56%对话会暴露秘密。

PrivacyBench: A Conversational Benchmark for Evaluating Privacy in Personalized AI

  • 构建含隐藏秘密的对话数据集,评估AI在多轮对话中保护隐私能力
  • 检索增强生成模型在26.56%交互中泄露秘密,提示优化仅降至5.12%
  • 指出当前架构依赖生成器单点防护,需系统性隐私设计

个性化AI代理依赖用户数字足迹,常涉及私密邮件、聊天记录和购物历史。但缺乏社会语境感知的系统可能无意泄露用户秘密,威胁数字福祉。我们提出PrivacyBench,一个基于社会情境的数据集,包含嵌入式秘密,并通过多轮对话评估秘密保护能力。测试检索增强生成(RAG)助手发现,其在高达26.56%的交互中泄露秘密。使用隐私意识提示可将泄漏降至5.12%,但仅部分缓解。检索机制仍无差别访问敏感数据,使隐私保护责任全压在生成器上,形成单点故障,当前架构不适合大规模部署。研究强调亟需结构性隐私优先设计,以保障网络的伦理与包容性。

原文摘要 · Abstract (English)

Personalized AI agents rely on access to a user's digital footprint, which often includes sensitive data from private emails, chats and purchase histories. Yet this access creates a fundamental societal and privacy risk: systems lacking social-context awareness can unintentionally expose user secrets, threatening digital well-being. We introduce PrivacyBench, a benchmark with socially grounded datasets containing embedded secrets and a multi-turn conversational evaluation to measure secret preservation. Testing Retrieval-Augmented Generation (RAG) assistants reveals that they leak secrets in up to 26.56% of interactions. A privacy-aware prompt lowers leakage to 5.12%, yet this measure offers only partial mitigation. The retrieval mechanism continues to access sensitive data indiscriminately, which shifts the entire burden of privacy preservation onto the generator. This creates a single point of failure, rendering current architectures unsafe for wide-scale deployment. Our findings underscore the urgent need for structural, privacy-by-design safeguards to ensure an ethical and inclusive web for everyone.

隐私安全AI评测RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。