arXiv:2605.03228cs.CRcs.AI2026-05被引 5

用影子记忆防范大模型代理的长期攻击威胁

MAGE: Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory

论文配图:MAGE: Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory
图 1 · 摘自论文原文
  • 构建安全专用的影子记忆,追踪全程关键上下文
  • 多数攻击可在早期阶段被准确检测,误报率低
  • 适合部署在金融、医疗等高风险场景的智能代理

随着大语言模型驱动的智能体被广泛用于复杂现实任务,其面临越来越多利用长时间交互实现恶意目标的攻击。这些长周期威胁对关键领域中的安全部署构成重大风险。本文提出MAGE(Memory As Guardrail Enforcement),一种新型防御框架,通过借鉴系统安全中的‘影子栈’思想,维护一个专注安全的代理记忆,持续提炼并保留执行全过程中的安全关键上下文,并利用该影子记忆在动作执行前主动评估风险。大量实验表明,MAGE在多种长周期威胁中显著优于现有防御方法,检测准确率高,多数攻击可在早期发现,且对代理性能影响极小。据我们所知,MAGE是首个基于代理记忆机制检测和缓解长周期威胁的框架,为该关键挑战建立了新范式,并开启了未来研究的新方向。

原文摘要 · Abstract (English)

As large language model (LLM)-powered agents are increasingly deployed to perform complex, real-world tasks, they face a growing class of attacks that exploit extended user-agent-environment interactions to pursue malicious objectives improbable in single-turn settings. Such long-horizon threats pose significant risks to the safe deployment of LLM agents in critical domains. In this paper, we present MAGE (Memory As Guardrail Enforcement), a novel defensive framework designed to counter a wide range of long-horizon threats. Inspired by the "shadow stack" abstraction in systems security, MAGE maintains a dedicated, safety-focused agentic memory that distills and retains safety-critical context across the agent's full execution trajectory, leveraging this shadow memory to proactively assess the risk of pending actions prior to their execution. Extensive evaluation demonstrates that MAGE substantially outperforms existing defenses across diverse long-horizon threats in detection accuracy, achieves early-stage detection for the majority of attacks, and introduces only negligible overhead to agent utility. To our best knowledge, MAGE represents the first framework to detect and mitigate long-horizon threats using an agentic memory approach, establishing a new paradigm for this critical challenge and opening promising directions for future research.

大模型安全智能体防御影子记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。