提出新框架区分用户偏好真实变化与恶意记忆污染,提升个性化大模型安全性。
CAPTURE: Disentangling Preference Drift from Memory Poisoning in Personalized LLM Agents
- 用神经微分方程追踪用户隐状态,多时标记忆账本记录行为演化
- 在96名用户测试中胜率71.5%,对真实偏好更新接受率达83.5%
- 适合关注个性化系统安全与鲁棒性的研究人员
个性化语言代理依赖持久记忆适应用户,但该机制也带来攻击风险。当新信息与存储偏好冲突时,代理需区分真实偏好漂移、临时上下文变化、模糊性或对抗性记忆污染。本文将问题建模为连续时间部分可观测决策过程,揭示仅依赖时效性和来源的规则不足。CAPTURE采用神经微分方程信念追踪器、多时标记忆账本、不确定性触发澄清及引用记忆的反事实审计。在96名用户的480个保留回合中,胜率达71.5%,高于同监督基线(69.3%)和最强启发式基线(66.1%)。固定策略攻击成功率降至11.5%,同时接受83.5%的真实偏好更新。面对可访问权重的自适应攻击者,攻击成功率升至24.7%,揭示了适应性与安全性之间的实际权衡。零样本评估在独立构建基准上表现稳健,并复现了40名用户持续2-3周的纵向交互历史。结果表明,显式建模偏好真实性可同时提升个性化与鲁棒性。
原文摘要 · Abstract (English)
Personalized language agents use persistent memory to adapt to users over time, but the same mechanism creates an attack surface. When new information conflicts with stored preferences, an agent must distinguish genuine preference drift from temporary context shifts, ambiguity, or adversarial memory poisoning. We formulate this problem as a continuous-time partially observable decision process over a latent user state and show why rules based only on recency and provenance are insufficient. CAPTURE addresses this ambiguity with a neural differential-equation belief tracker, a multi-timescale memory ledger, uncertainty-triggered clarification, and counterfactual auditing of cited memories. On 480 held-out episodes from 96 users, CAPTURE achieves a 71.5% win rate, compared with 69.3% for an identically supervised baseline and 66.1% for the strongest heuristic baseline. It limits fixed-policy poisoning success to 11.5% while accepting 83.5% of genuine preference updates. Under an adaptive attacker with access to the released weights, attack success rises to 24.7%, exposing a real adaptation-security tradeoff. We further evaluate the frozen system zero-shot on an independently constructed benchmark and replay longitudinal interaction histories from 40 users collected over two to three weeks. These results suggest that modeling preference authenticity explicitly can improve both personalization and robustness in memory-augmented LLM agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。