arXiv:2606.04067cs.CRcs.AI2026-06被引 1

基于上下文完整性,智能删减大模型请求中的敏感信息。

Need to Know: Contextual-Integrity-Grounded Query Rewriting for Privacy-Conscious LLM Delegation

论文配图:Need to Know: Contextual-Integrity-Grounded Query Rewriting for Privacy-Conscious LLM Delegation
图 1 · 摘自论文原文
  • 根据任务必要性判断敏感内容是否该保留
  • 在3167个样本上训练,隐私保护下提升10.1%任务效率
  • 适合需要保护隐私的云端大模型应用开发者

随着大模型融入日常使用,用户向云端模型提交的查询常混合任务必需内容与非必需敏感信息。现有基于类型的个人身份信息(PII)删除方法缺乏上下文感知,可能导致过度披露或误删关键回答内容。本文提出基于上下文完整性的隐私保护查询重写框架:仅当片段对任务必要时才转发。构建首个面向隐私委托的任务型上下文完整性基准 DelegateCI-Bench,包含3,167个样本,涵盖11项任务、20种任务类型,结合高质量合成数据、真实用户对话(WildChat)及高密度敏感信息的医疗挑战集。基于此基准,设计一种受上下文完整性引导的强化学习框架,将必要与非必要敏感片段转化为可验证优化信号,训练查询重写器在保留任务关键信息的同时抑制不必要的敏感披露。实验表明,所提学习重写器在隐私-效用权衡上表现最佳,平均效用相比本地基线最高提升10.1%。

原文摘要 · Abstract (English)

As LLMs become increasingly woven into everyday workflows, user queries sent to cloud hosted LLMs routinely mix task-essential content with task non-essential sensitive disclosures, yet type based PII redaction is context agnostic and may raise two issues: over disclosing untyped sensitive context and over removing answer bearing spans. We recast privacy preserving query rewriting under Contextual Integrity: a span should be forwarded only if it is necessary for the task. We introduce DelegateCI-Bench, the first task based Contextual Integrity benchmark for privacy-conscious delegation, comprising 3,167 samples that combine high quality synthetic data spanning 11 tasks and 20 task types, WildChat based real user queries, and a medical challenge set with dense sensitive information. Building on this benchmark, we propose a CI-guided reinforcement learning framework that converts essential and non-essential sensitive spans into verifiable optimization signals, and train a query rewriter to preserve task critical information while suppressing unnecessary sensitive disclosure. Experiments show that our learned rewriter achieves the best privacy-utility tradeoff, achieving up to +10.1 average utility over on-device baselines.

隐私保护大模型查询重写强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。