arXiv:2506.01055cs.CRcs.CL2025-06被引 33

攻击者可诱导智能代理泄露执行任务时观察到的个人数据。

Simple Prompt Injection Attacks Can Leak Personal Data Observed by LLM Agents During Task Execution

  • 设计基于数据流的注入攻击,触发代理泄露信息
  • 平均攻击成功率约15%-20%,部分任务达50个百分点性能下降
  • 涉及数据提取的任务最易被攻破,适合安全研究者参考

以往针对大语言模型的提示注入评估多聚焦于通用任务与攻击,难以揭示数据外泄等复杂威胁。本文研究提示注入如何导致工具调用型代理在任务执行中泄露个人数据。以虚构银行代理为例,构建基于数据流的攻击,并集成至近期的代理安全基准AgentDojo。为扩展评估范围,还创建了一个更丰富的合成人类-AI银行对话数据集。在AgentDojo的16个用户任务中,模型性能平均下降15-50个百分点,平均攻击成功率(ASR)约为20%;部分防御措施可将ASR降至零。尽管多数模型在被诱骗后仍避免泄露密码等高敏感数据(可能因安全对齐),但其他个人信息仍存在泄露风险。当密码与其他一两个个人信息一同被请求时,泄露概率上升。在48项任务的扩展评估中,平均ASR约为15%,且无内置防御能完全阻止泄露。涉及数据提取或授权流程的任务表现出最高攻击成功率,凸显任务类型、代理表现与防御效果间的交互关系。

原文摘要 · Abstract (English)

Previous benchmarks on prompt injection in large language models (LLMs) have primarily focused on generic tasks and attacks, offering limited insights into more complex threats like data exfiltration. This paper examines how prompt injection can cause tool-calling agents to leak personal data observed during task execution. Using a fictitious banking agent, we develop data flow-based attacks and integrate them into AgentDojo, a recent benchmark for agentic security. To enhance its scope, we also create a richer synthetic dataset of human-AI banking conversations. In 16 user tasks from AgentDojo, LLMs show a 15-50 percentage point drop in utility under attack, with average attack success rates (ASR) around 20 percent; some defenses reduce ASR to zero. Most LLMs, even when successfully tricked by the attack, avoid leaking highly sensitive data like passwords, likely due to safety alignments, but they remain vulnerable to disclosing other personal data. The likelihood of password leakage increases when a password is requested along with one or two additional personal details. In an extended evaluation across 48 tasks, the average ASR is around 15 percent, with no built-in AgentDojo defense fully preventing leakage. Tasks involving data extraction or authorization workflows, which closely resemble the structure of exfiltration attacks, exhibit the highest ASRs, highlighting the interaction between task type, agent performance, and defense efficacy.

提示注入数据泄露智能代理安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。