arXiv:2509.23614cs.AI2025-09被引 10

为大模型代理定制个性化安全防护,动态跟踪风险累积。

PSG-Agent: Personality-Aware Safety Guardrail for LLM-based Agents

  • 基于用户历史与实时输入生成个性化风险阈值
  • 跨轮次监控多环节风险,提升安全检测连续性
  • 适用于医疗、金融等高风险场景的智能防护

大模型代理在关键应用中的安全部署亟需有效防护机制。现有防护方法存在两大缺陷:一是对所有用户采用统一策略,忽视不同用户对相同行为的风险感知差异;二是逐条检查响应,忽略风险在多轮交互中的演化与累积。为此,我们提出PSG-Agent,一种个性化的动态防护系统。首先,通过挖掘交互历史中的稳定特征并结合当前查询的实时状态,生成用户特定的风险阈值与保护策略。其次,通过计划监控器、工具防火墙、响应守护、记忆守护等专用组件,在代理全流程中实现持续风险监测,追踪跨轮次风险积累并输出可验证的判断结果。最后,在医疗、金融及日常自动化等多样化场景下验证了PSG-Agent的有效性,其显著优于LlamaGuard3和AGrail等现有方案,为大模型代理提供了可执行、可审计的个性化安全保障路径。

原文摘要 · Abstract (English)

Effective guardrails are essential for safely deploying LLM-based agents in critical applications. Despite recent advances, existing guardrails suffer from two fundamental limitations: (i) they apply uniform guardrail policies to all users, ignoring that the same agent behavior can harm some users while being safe for others; (ii) they check each response in isolation, missing how risks evolve and accumulate across multiple interactions. To solve these issues, we propose PSG-Agent, a personalized and dynamic system for LLM-based agents. First, PSG-Agent creates personalized guardrails by mining the interaction history for stable traits and capturing real-time states from current queries, generating user-specific risk thresholds and protection strategies. Second, PSG-Agent implements continuous monitoring across the agent pipeline with specialized guards, including Plan Monitor, Tool Firewall, Response Guard, Memory Guardian, that track cross-turn risk accumulation and issue verifiable verdicts. Finally, we validate PSG-Agent in multiple scenarios including healthcare, finance, and daily life automation scenarios with diverse user profiles. It significantly outperform existing agent guardrails including LlamaGuard3 and AGrail, providing an executable and auditable path toward personalized safety for LLM-based agents.

大模型安全个性化防护动态监控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。