arXiv:2601.14660cs.CRcs.AI2026-01

用激活值检测保护私密信息,让聊天机器人更懂边界。

NeuroFilter: Activation-Based Guardrails for Privacy-Conscious LLM Agents

  • 通过分析模型内部激活值,动态识别敏感信息
  • 在单轮和多轮对话中均有效防止隐私泄露
  • 无需额外模型,适合实时部署的隐私防护

智能大语言模型(LLM)具备推理、规划和操作非结构化数据的能力,正推动个人助理、金融与法律等领域变革。但其高效运作往往需访问敏感个人信息,带来严重的运行时隐私风险,尤其在上下文相关的信息披露方面。现有研究指出,这类模型难以稳定遵守隐私规范;而现有的防御手段多依赖额外的辅助大模型监控,成本高且对语义屏蔽攻击缺乏鲁棒性。本文提出基于激活值探测的隐私过滤机制,首次系统性地探究了对话轨迹中模型内部状态的演化过程,突破了静态单次提示分析的局限。实验表明,该方法在单轮与多轮对话场景下均兼具计算高效性和强有效性,为构建隐私敏感型智能体提供了新路径。

原文摘要 · Abstract (English)

Agentic Large Language Models (LLMs) are models able to reason, plan, and execute tools over unstructured data. These abilities are enabling transformative applications in domains spanning from personal assistant, financial, and legal domains. While these systems can substantially improve productivity and service quality, effective agency typically requires access to sensitive personal or organizational information. However, this access introduces critical inference-time privacy risks, specifically regarding contextually appropriate information disclosure. While recent studies highlight the inability of agentic LLMs to consistently adhere to privacy norms, existing defenses often rely on auxiliary LLM-based monitors. However, these defenses are expensive and offer limited protection against attacks that are robust to semantic censorship. To contrast this background, this paper proposes a notion of privacy filters based on activation probing. We show that these filters are both computationally efficient and effective for both single-turn and multi-turn conversational settings. Furthermore, this work provides the first systematic investigation into probing model internals across a conversation trajectory, moving beyond static, single-prompt analysis to capture the evolving state of privacy-sensitive interactions.

隐私保护大模型激活值智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。