arXiv:2604.00209cs.CL2026-04被引 2

发现大模型内部有隐私规范结构,可精准调控避免泄露。

Do LLMs Know What Is Private Internally? Probing and Steering Contextual Privacy Norms in Large Language Model Representations

论文配图:Do LLMs Know What Is Private Internally? Probing and Steering Contextual Privacy Norms in Large Language Model Representations
图 1 · 摘自论文原文
  • 用上下文完整性理论探测模型内部隐私规则
  • 三类隐私参数在激活空间线性独立,可分别控制
  • 新方法比传统方式更有效减少隐私泄露,适合安全开发

大型语言模型在高风险场景中部署日益增多,却常在人类会谨慎处理的情况下披露私密信息。这引出一个根本问题:大模型是否内在编码了情境隐私规范?若有,为何仍频繁违规?我们首次系统研究了基于情境完整性(CI)理论的模型内部隐私规范结构。通过探测多个模型,发现决定隐私规范的三个核心参数——信息类型、接收方与传播原则——在激活空间中表现为线性可分且功能独立的方向。尽管存在这种内部结构,模型实际仍会泄露私密信息,暴露出概念表征与行为之间的明显差距。为此,我们提出基于CI参数的定向调控方法,可独立干预每个维度。该结构化控制方式比整体式干预更有效、更可预测地减少隐私违规。结果表明,情境隐私失败源于表征与行为间的错位,而非缺乏意识;利用CI的组合结构可实现更可靠的隐私控制,为提升大模型的情境隐私理解提供新路径。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed in high-stakes settings, yet they frequently violate contextual privacy by disclosing private information in situations where humans would exercise discretion. This raises a fundamental question: do LLMs internally encode contextual privacy norms, and if so, why do violations persist? We present the first systematic study of contextual privacy as a structured latent representation in LLMs, grounded in contextual integrity (CI) theory. Probing multiple models, we find that the three norm-determining CI parameters (information type, recipient, and transmission principle) are encoded as linearly separable and functionally independent directions in activation space. Despite this internal structure, models still leak private information in practice, revealing a clear gap between concept representation and model behavior. To bridge this gap, we introduce CI-parametric steering, which independently intervenes along each CI dimension. This structured control reduces privacy violations more effectively and predictably than monolithic steering. Our results demonstrate that contextual privacy failures arise from misalignment between representation and behavior rather than missing awareness, and that leveraging the compositional structure of CI enables more reliable contextual privacy control, shedding light on potential improvement of contextual privacy understanding in LLMs.

隐私保护模型可控性大模型机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。