首个评估大模型记忆演进安全风险的基准,揭示记忆污染导致行为失控。
MemEvoBench: Benchmarking Safety Risks from Memory Misevolution in LLM Agents

- 构建多轮交互下的记忆演化测试框架,模拟误导信息注入
- 36种风险类型下,模型安全性能显著下降,尤其在有偏记忆更新时
- 提醒开发者关注记忆演化安全,静态防御手段已不足
赋予大型语言模型持久记忆可提升交互连续性和个性化体验,但也会引入新的安全风险。例如,污染或有偏记忆的累积可能引发异常代理行为。现有评估方法尚未建立针对记忆演进问题的标准化框架。本文提出MemEvoBench,首个评估大模型代理在长期记忆演化中安全性的问题基准。该框架包含7个领域、36种风险类型的问答式任务,以及基于20个Agent-SafetyBench环境改造的工作流任务,模拟工具输出噪声与有偏反馈。两类任务均在多轮交互中混合良性与误导性记忆池,以模拟记忆演化过程。对代表性模型的实验表明,在有偏记忆更新下,安全性能显著下降。分析显示,记忆演化是此类失败的重要诱因。此外,基于静态提示的防御策略效果有限,凸显保障记忆演化安全的紧迫性。
原文摘要 · Abstract (English)
Equipping Large Language Models (LLMs) with persistent memory enhances interaction continuity and personalization but introduces new safety risks. Specifically, contaminated or biased memory accumulation can trigger abnormal agent behaviors. Existing evaluation methods have not yet established a standardized framework for measuring memory misevolution. This phenomenon refers to the gradual behavioral drift resulting from repeated exposure to misleading information. To address this gap, we introduce MemEvoBench, the first benchmark evaluating long-horizon memory safety in LLM agents against adversarial memory injection, noisy tool outputs, and biased feedback. The framework consists of QA-style tasks across 7 domains and 36 risk types, complemented by workflow-style tasks adapted from 20 Agent-SafetyBench environments with noisy tool returns. Both settings employ mixed benign and misleading memory pools within multi-round interactions to simulate memory evolution. Experiments on representative models reveal substantial safety degradation under biased memory updates. Our analysis suggests that memory evolution is a significant contributor to these failures. Furthermore, static prompt-based defenses prove insufficient, underscoring the urgency of securing memory evolution in LLM agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。