arXiv:2508.20282cs.CRcs.AI2025-08被引 5

本地部署的智能研究代理会泄露用户隐私,因网络行为有独特模式可被追踪。

Network-Level Prompt and Trait Leakage in Local Research Agents

  • 利用访问域名和时间戳等网络元数据,推断用户提问内容和隐性特征。
  • 攻击可恢复73%以上提示功能信息,多会话下准确识别19个潜在用户属性。
  • 适用于关注隐私安全的研究者、企业及政策制定者,尤其在本地部署场景中。

我们发现,基于语言模型的网页与研究代理(WRAs)——用于在互联网上探索复杂话题的系统——容易受到被动网络观察者的推理攻击。组织和个人在本地部署WRAs以保障隐私、合规或节省成本时,会暴露于DNS解析器、恶意ISP、VPN、网页代理、企业或政府防火墙之下。不同于人类零星稀疏的浏览行为,WRAs每次请求会访问70至140个不同域名,并具有独特的时间模式,带来显著隐私风险。本文提出一种新型的提示与用户特征泄露攻击,仅依赖网络级元数据(即访问的IP地址及其时间)。我们构建了基于真实用户查询与合成人格生成查询的新数据集,并引入行为度量指标OBELS,全面评估原始提示与推断结果的相似性。结果显示,该攻击能恢复超过73%的提示功能与领域知识。在多会话设置下,可高精度恢复32个潜在特征中的19个。即使在部分可观测和噪声环境下,攻击依然有效。最后,我们讨论了缓解策略,如限制域名多样性或混淆访问痕迹,可在几乎不影响使用效果的前提下,平均降低29%的攻击成功率。

原文摘要 · Abstract (English)

We show that Web and Research Agents (WRAs) -- language-model-based systems that investigate complex topics on the Internet -- are vulnerable to inference attacks by passive network observers. Deployment of WRAs \emph{locally} by organizations and individuals for privacy, legal, or financial purposes exposes them to DNS resolvers, malicious ISPs, VPNs, web proxies, and corporate or government firewalls. However, unlike sporadic and scarce web browsing by humans, WRAs visit $70{-}140$ domains per each request with a distinct timing pattern creating unique privacy risks. Specifically, we demonstrate a novel prompt and user trait leakage attack against WRAs that only leverages their network-level metadata (i.e., visited IP addresses and their timings). We start by building a new dataset of WRA traces based on real user search queries and queries generated by synthetic personas. We define a behavioral metric (called OBELS) to comprehensively assess similarity between original and inferred prompts, showing that our attack recovers over 73\% of the functional and domain knowledge of user prompts. Extending to a multi-session setting, we recover up to 19 of 32 latent traits with high accuracy. Our attack remains effective under partial observability and noisy conditions. Finally, we discuss mitigation strategies that constrain domain diversity or obfuscate traces, showing negligible utility impact while reducing attack effectiveness by an average of 29\%.

隐私安全网络侦查大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。