arXiv:2605.27825cs.CRcs.LG2026-05被引 1

针对聊天代理记忆的隐私攻击,可有效识别敏感信息是否被存储。

MRMMIA: Membership Inference Attacks on Memory in Chat Agents

论文配图:MRMMIA: Membership Inference Attacks on Memory in Chat Agents
图 1 · 摘自论文原文
  • 通过多轮召回探测提取记忆归属信号,适配黑盒灰盒白盒场景。
  • 在多种设置下均超越基线方法,证明记忆存在显著隐私泄露风险。
  • 为聊天代理记忆隐私评估提供首个系统性框架,适合安全研究者参考。

成员推理攻击(MIAs)用于检测目标数据记录是否属于系统的私有数据,已成为衡量机器学习系统隐私泄露的标准工具。以往研究主要聚焦于训练数据集或检索数据库,但针对代理记忆的攻击关注较少,尽管此类记忆可能包含敏感的用户-代理交互、检索事实及用户偏好。本文聚焦聊天代理记忆的成员推理攻击,提出多召回记忆成员推理攻击(MRMMIA),一种统一框架,利用多个召回探测在黑盒、灰盒和白盒环境下提取成员信号。实验表明,MRMMIA始终优于基线方法。结果揭示了聊天代理记忆中的隐私风险,并为该类系统的成员泄露提供了初步评估框架。

原文摘要 · Abstract (English)

Membership inference attacks (MIAs) test whether a target data record belongs to a system's private data, and have become a standard tool to measure privacy leakage in machine learning systems. Prior work has primarily focused on training corpora or retrieval databases. However, MIAs against agent memory have received less attention, even though such memory can contain sensitive user-agent interactions, retrieved facts, and user preferences. Therefore, in this work, we focus on chat agent memory MIAs, where an adversary infers whether a candidate memory unit belongs to the chat agent's memory store. We propose Multi-Recall Memory MIA (MRMMIA), a unified attack that utilizes multiple recall probes to the agent to extract the membership signal across black-box, gray-box, and white-box settings. Our experiments demonstrate that MRMMIA consistently outperforms baselines. Our results expose the privacy risk in agents and provide an initial evaluation framework for membership leakage in chat-agent memory systems.

隐私攻击聊天代理成员推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。