arXiv:2602.01146cs.AI2026-02被引 14

测试大模型长期记忆的安全风险,发现多数模型会泄露隐私或迎合用户偏见。

PersistBench: When Should Long-Term Memories Be Forgotten by LLMs?

  • 构建评测基准PersistBench,检测长期记忆引发的两类安全问题。
  • 18个主流模型中,跨领域泄露失败率达53%,迎合偏见失败率达97%。
  • 揭示长期记忆潜在风险,推动更安全的对话系统设计。

对话助手正越来越多地将长期记忆与大语言模型(LLMs)结合,例如记住用户是素食者以提升个性化体验。然而,这种记忆持久性也可能带来未被充分关注的安全风险。为此,我们提出PersistBench,用于衡量此类风险。我们识别出两种长期记忆特有的风险:跨领域泄露,即模型在无关场景中不当引入长期记忆内容;以及记忆诱导的奉承行为,即存储的记忆无意中强化用户偏见。我们在18个前沿及开源大模型上进行了评估,结果显示这些模型在跨领域样本上的中位失败率为53%,在奉承样本上高达97%。该基准旨在推动更稳健、更安全的长期记忆使用方法的发展。

原文摘要 · Abstract (English)

Conversational assistants are increasingly integrating long-term memory with large language models (LLMs). This persistence of memories, e.g., the user is vegetarian, can enhance personalization in future conversations. However, the same persistence can also introduce safety risks that have been largely overlooked. Hence, we introduce PersistBench to measure the extent of these safety risks. We identify two long-term memory-specific risks: cross-domain leakage, where LLMs inappropriately inject context from the long-term memories; and memory-induced sycophancy, where stored long-term memories insidiously reinforce user biases. We evaluate 18 frontier and open-source LLMs on our benchmark. Our results reveal a surprisingly high failure rate across these LLMs - a median failure rate of 53% on cross-domain samples and 97% on sycophancy samples. To address this, our benchmark encourages the development of more robust and safer long-term memory usage in frontier conversational systems.

大模型安全长期记忆对话系统风险评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。