arXiv:2605.18133cs.CRcs.AI2026-05

黑盒聊天机器人中,攻击者可借伪造内容诱导隐私泄露

An Empirical Study of Privacy Leakage Chains via Prompt Injection in Black-Box Chatbot Environments

  • 用外部内容伪装成良性数据,诱导代理执行恶意任务
  • 新方法'示例化'攻击成功率高于旧的假完成技术
  • 实验证明可构建隐私数据外泄链,适合安全研究者关注

基于大模型的聊天机器人代理通过结合自然语言推理与外部工具(如网页浏览)来处理用户请求,提升了可用性,但也带来了安全隐患。本文研究了在无模型权重、系统提示或代理实现细节的黑盒环境下,通过间接提示注入引发的隐私泄露攻击链。首先分析攻击者如何构造看似无害的外部内容,劫持代理原定任务;随后评估一种新提示注入技术——示例化,利用外部内容中的桥梁将用户提示和检索页面开头重构成少样本示例,再附加攻击目标;对比显示其攻击成功率高于先前的假完成技术。最后,在受控环境中演示了基于虚构个人信息的数据外泄链。结果表明,提示注入、越狱式指令操控与网络工具调用可组合成实际可行的隐私泄露路径。

原文摘要 · Abstract (English)

LLM-based chatbot agents increasingly process user requests by combining natural-language reasoning with external tools such as web browsing. These capabilities improve usability, but they also create attack surfaces when untrusted external content is processed as part of a user' s task. This paper studies a privacy-leakage attack chain based on indirect prompt injection in black-box chatbot environments, where the attacker has no access to model weights, system prompts, or agent implementation details including how a trajectory is actually managed during its processing for a query. We first analyze how an attacker can hijack an agent' s intended task by crafting external content that appears benign to the victim while inducing the agent to execute an attacker-defined objective. We then evaluate a new prompt-injection technique, called exemplification, which uses a bridge in the external content to reframe the user prompt and the benign beginning of the retrieved page as few-shot examples before appending the attacker' s objective. We compare its attack success rate with a prior fake-completion technique. Finally, we demonstrate a proof-of-concept data-exfiltration chain using fictitious personal information in a controlled setting. Our results suggest that prompt injection, jailbreak-style instruction steering, and web-tool invocation can be combined into a feasible privacy-leakage path in deployed chatbot agents.

隐私泄露提示注入黑盒攻击大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。