让大模型代理在执行中动态决策隐私保护,提升安全与有用性平衡。
Contextualized Privacy Defense for LLM Agents
- 用指导模型生成每一步的隐私建议,主动引导行为。
- 隐私保护率达94.2%,帮助性达80.6%,优于现有方法。
- 适合需高隐私安全的多步任务代理系统使用。
大语言模型代理越来越多地处理用户个人信息,但现有隐私防护手段在设计和适应性上仍显不足。多数先前方法依赖静态或被动防御,如提示词控制和限制机制,难以支持多步骤执行中的上下文感知、主动决策。本文提出情境化防御指导(CDI),一种新型隐私防护范式:在执行过程中,由指导模型生成针对具体步骤的、上下文敏感的隐私指引,主动塑造代理行为,而非仅约束或否决。关键在于,CDI 配合基于经验的优化框架,通过强化学习训练指导模型,将发生隐私违规的失败轨迹转化为学习环境。我们将基线防御与CDI形式化为标准代理循环中的不同干预点,并在统一仿真框架中比较其隐私保护与帮助性的权衡。结果表明,CDI 在隐私保护率(94.2%)与帮助性(80.6%)之间保持更优平衡,且对对抗条件更具鲁棒性,泛化能力更强。
原文摘要 · Abstract (English)
LLM agents increasingly act on users' personal information, yet existing privacy defenses remain limited in both design and adaptability. Most prior approaches rely on static or passive defenses, such as prompting and guarding. These paradigms are insufficient for supporting contextual, proactive privacy decisions in multi-step agent execution. We propose Contextualized Defense Instructing (CDI), a new privacy defense paradigm in which an instructor model generates step-specific, context-aware privacy guidance during execution, proactively shaping actions rather than merely constraining or vetoing them. Crucially, CDI is paired with an experience-driven optimization framework that trains the instructor via reinforcement learning (RL), where we convert failure trajectories with privacy violations into learning environments. We formalize baseline defenses and CDI as distinct intervention points in a canonical agent loop, and compare their privacy-helpfulness trade-offs within a unified simulation framework. Results show that our CDI consistently achieves a better balance between privacy preservation (94.2%) and helpfulness (80.6%) than baselines, with superior robustness to adversarial conditions and generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。