心跳机制让AI悄悄中毒,污染记忆影响用户行为
Mind Your HEARTBEAT! Claw Background Execution Inherently Enables Silent Memory Pollution
- 背景心跳执行共享会话内存,外部信息直接进入对话上下文
- 61%误导率、91%长期记忆固化、76%跨会话影响已发生
- 无需提示注入,普通社交假信息即可无声污染AI行为
我们发现主流Claw个人AI代理存在关键安全漏洞:在心跳驱动的后台执行中遇到的不可信内容会无声污染代理内存,并间接影响用户交互行为而用户毫无察觉。该漏洞源于整个Claw生态共有的架构设计——心跳后台执行与用户前台对话共享同一会话,导致从邮件、消息通道、新闻源、代码库及社交平台等外部来源摄入的内容可进入与前端交互相同的内存上下文,且用户可见性有限、来源溯源不清。我们将其建模为暴露(E)→记忆(M)→行为(B)路径:心跳期间接触的虚假信息进入短期会话上下文,可能被写入长期记忆,进而塑造后续用户行为。通过在代理内社交场景中构建受控研究副本MissClaw验证发现:(1)社会可信度线索(尤其是感知共识)是短期行为影响的主要驱动力,误导率最高达61%;(2)常规内存保存行为可将短期污染转化为持久长期记忆,转化率高达91%,跨会话行为影响达76%;(3)在自然浏览情境下,尽管存在内容稀释与上下文修剪,污染仍能跨越会话边界。总体而言,无需提示注入,日常社交假信息即可在心跳驱动后台执行下悄然塑造代理记忆与行为。
原文摘要 · Abstract (English)
We identify a critical security vulnerability in mainstream Claw personal AI agents: untrusted content encountered during heartbeat-driven background execution can silently pollute agent memory and subsequently influence user-facing behavior without the user's awareness. This vulnerability arises from an architectural design shared across the Claw ecosystem: heartbeat background execution runs in the same session as user-facing conversation, so content ingested from any external source monitored in the background (including email, message channels, news feeds, code repositories, and social platforms) can enter the same memory context used for foreground interaction, often with limited user visibility and without clear source provenance. We formalize this process as an Exposure (E) $\rightarrow$ Memory (M) $\rightarrow$ Behavior (B) pathway: misinformation encountered during heartbeat execution enters the agent's short-term session context, potentially gets written into long-term memory, and later shapes downstream user-facing behavior. We instantiate this pathway in an agent-native social setting using MissClaw, a controlled research replica of Moltbook. We find that (1) social credibility cues, especially perceived consensus, are the dominant driver of short-term behavioral influence, with misleading rates up to 61%; (2) routine memory-saving behavior can promote short-term pollution into durable long-term memory at rates up to 91%, with cross-session behavioral influence reaching 76%; (3) under naturalistic browsing with content dilution and context pruning, pollution still crosses session boundaries. Overall, prompt injection is not required: ordinary social misinformation is sufficient to silently shape agent memory and behavior under heartbeat-driven background execution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。