提出用替代表达保护隐私,比删减或抽象更有效。
Say Something Else: Rethinking Contextual Privacy as Information Sufficiency
- 引入自由文本伪名化,用等效替代敏感信息
- 多轮对话中伪名化隐私保护效果最优,泄露率低16.3%
- 适合关注隐私安全的AI交互设计者
大语言模型代理越来越多地代用户起草消息,但用户常过度披露敏感信息且对隐私定义不一致。现有系统仅支持屏蔽(删除)和泛化(替换为抽象词),且通常在孤立消息上评估,策略空间与评估环境均不完整。本文将隐私保护通信形式化为信息充分性(IS)任务,提出自由文本伪名化作为第三种策略,用功能等价的替代项替换敏感属性,并设计对话式评估协议,在真实多轮追问压力下测试策略。我们在792个场景中评估七款前沿LLM,覆盖三种权力关系类型(机构、同侪、亲密)和三种敏感度类别(歧视风险、社交成本、边界)。结果表明,伪名化在隐私-效用权衡上表现最佳;单消息评估系统性低估泄露,泛化在多轮追问下隐私损失最高达16.3个百分点。
原文摘要 · Abstract (English)
LLM agents increasingly draft messages on behalf of users, yet users routinely overshare sensitive information and disagree on what counts as private. Existing systems support only suppression (omitting sensitive information) and generalization (replacing information with an abstraction), and are typically evaluated on single isolated messages, leaving both the strategy space and evaluation setting incomplete. We formalize privacy-preserving LLM communication as an \textbf{Information Sufficiency (IS)} task, introduce \textbf{free-text pseudonymization} as a third strategy that replaces sensitive attributes with functionally equivalent alternatives, and propose a \textbf{conversational evaluation protocol} that assesses strategies under realistic multi-turn follow-up pressure. Across 792 scenarios spanning three power-relation types (institutional, peer, intimate) and three sensitivity categories (discrimination risk, social cost, boundary), we evaluate seven frontier LLMs on privacy at two granularities, covertness, and utility. Pseudonymization yields the strongest privacy\textendash utility tradeoff overall, and single-message evaluation systematically underestimates leakage, with generalization losing up to 16.3 percentage points of privacy under follow-up.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。