arXiv:2605.18930cs.CRcs.AI2026-05被引 5

攻击者用看似合理但有害的体验,让大模型自我进化时形成错误规则。

OEP: Poisoning Self-Evolving LLM Agents via Locally Correct but Non-Transferable Experiences

论文配图:OEP: Poisoning Self-Evolving LLM Agents via Locally Correct but Non-Transferable Experiences
图 1 · 摘自论文原文
  • 构造局部正确但无法迁移的假经验,诱导模型错误反思
  • 在GPT-4o上实现超50%的攻击成功率,且能绕过安全检测
  • 适合研究模型安全与对抗性训练的学者关注

带记忆的大型语言模型(LLM)代理通过迭代反思和自我演化解决复杂任务,但这一机制带来安全风险。现有代理记忆攻击需特权访问或恶意内容,易被高级安全过滤器识别。我们发现,攻击者可诱导代理生成表面合理、语义可信却导致有害泛化的体验。基于此,提出低权限黑盒攻击方法OEP,无需直接控制系统提示或内存数据库。OEP构建包含局部正确解法、不可迁移方法及严重后果的对抗性边缘案例,使反思偏向规避风险的规则形成。在记忆整合阶段,代理可能过度信任自生成反思,将局部经验提炼为高优先级但过度泛化的规则,引发下游任务失败。三个领域评估显示,OEP在使用GPT-4o代理时攻击成功率超过50%,且在大模型审计防御下仍优于现有攻击。

原文摘要 · Abstract (English)

Memory-augmented large language model (LLM) agents use iterative reflection and self-evolution to solve complex tasks, but these mechanisms introduce security risks. Existing agentic memory attacks require privileged access or explicit malicious content, making them detectable by advanced safety filters. This leaves a subtler attack surface underexplored: whether adversaries can induce agent to generate experiences that appear locally correct and semantically plausible yet induce harmful generalization during reflection. We find that reflective agents are vulnerable to such clean experiences, especially when paired with severe but plausible hypothetical consequences. Based on this observation, we introduce Obsessive Experience Poisoning (OEP), a low-privilege black-box attack requiring no direct control over the system prompt or memory database. OEP constructs adversarial clean edge-cases that combine locally correct solutions, non-transferable methods, and severe consequences, biasing reflection toward risk-averse rule formation. During memory consolidation, agents may over-trust self-generated reflections and distill localized experiences into high-priority but over-generalized rules, causing downstream failures. Evaluations across three domains show that OEP achieves ASR above 50\% with GPT-4o agents, and outperforms existing attacks under LLM auditing defense.

模型安全对抗攻击自进化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。