arXiv:2608.21209cs.AIcs.CL2026-08

让大模型按用户个人偏好控制隐私披露,提升个性化保护能力。

Personalized Privacy Control in LLMs via Attention Head Intervention

论文配图:Personalized Privacy Control in LLMs via Attention Head Intervention
图 1 · 摘自论文原文
  • 通过注意力头干预,在推理阶段动态调整模型隐私行为。
  • 现有提示策略对个性化隐私政策遵守率不足52%(Qwen2.5-7B)。
  • 适合关注用户隐私差异、需精准控制数据泄露的研究者。

代理型AI使大模型能够访问多样化的用户数据,引发严峻的隐私问题。已有研究聚焦于上下文依赖的隐私规范,但同一上下文中不同用户的可接受披露边界可能不同。为此,本文提出「个性化隐私」概念,将用户特定的披露偏好纳入隐私控制机制。我们进一步构建了P3Bench(个性化隐私保留基准),在原有上下文隐私策略基础上扩展个性化披露策略。实验表明,基于提示的策略无法可靠执行个性化隐私策略:Qwen2.5-7B与Gemma3-4B平均忽略率达51.25%和74.28%。为解决此问题,我们提出 extsc{Repair},一种鲁棒的推理时注意力头干预方法,可有效调整模型输出以符合用户隐私政策。该方法显著提升模型对用户特定隐私偏好的遵循度,减少不合规响应情况。

原文摘要 · Abstract (English)

The rise of agentic AI enables LLMs to access diverse user data, raising critical privacy concerns. Prior work on contextual privacy studies whether LLMs regulate information disclosure according to context-dependent norms. However, acceptable disclosure boundaries may vary across users even within the same context. To address this limitation, we introduce \textit{personalized privacy}, which incorporates user-specific disclosure preferences into privacy control. We further present P3Bench~(\textbf{P}ersonalized \textbf{P}rivacy \textbf{P}reservation \textbf{Bench}mark), a novel benchmark extending contextual privacy policies with personalized disclosure policies. Experiments show that prompt-based policies fail to reliably enforce personalized privacy policies, with Qwen2.5-7B and Gemma3-4B showing average policy ignorance ratios of 51.25\% and 74.28\%, respectively. Finally, to address this problem, we propose \textsc{Repair}, a robust inference-time attention head intervention method that adjusts disclosure behavior toward policy-consistent responses. Our method significantly improves adherence to user-specific privacy preferences by reducing cases where the model fails to follow the given policy.

大模型隐私控制注意力干预个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。