评估智能体工作流中的隐私泄露风险,发现80%场景存在中间环节泄露。
AgentSCOPE: Evaluating Contextual Privacy Across Agentic Workflows
- 构建隐私流图,追踪每个信息环节的隐私边界。
- 在62个场景中,80%出现隐私违规,多数发生在工具响应阶段。
- 适合关注AI系统隐私安全的研究者与开发者参考。
智能体系统正越来越多地代表用户处理日常任务,访问日历、邮件和个人文件。现有隐私评估主要聚焦输入输出边界,但任务执行中涉及多个中间信息流(如智能体查询到工具响应),这些环节尚未被评估。我们认为,智能体流水线中的每一处边界都可能是隐私泄露点,需独立评估。为此,我们提出基于情境完整性的隐私流图框架,将智能体执行分解为一系列信息流,每条流标注五种情境完整性参数,并可追溯违规源头。我们构建了包含62个多工具场景的AgentSCOPE基准,覆盖八个监管领域,各阶段均提供真实标签。对七款先进LLM的评估显示,超过80%的场景存在隐私违规,即使最终输出看似无害(仅24%),大多数违规发生在工具响应阶段,此时API indiscriminately 返回敏感数据。结果表明,仅评估输出会严重低估智能体系统的隐私风险。
原文摘要 · Abstract (English)
Agentic systems are increasingly acting on users' behalf, accessing calendars, email, and personal files to complete everyday tasks. Privacy evaluation for these systems has focused on the input and output boundaries, but each task involves several intermediate information flows, from agent queries to tool responses, that are not currently evaluated. We argue that every boundary in an agentic pipeline is a site of potential privacy violation and must be assessed independently. To support this, we introduce the Privacy Flow Graph, a Contextual Integrity-grounded framework that decomposes agentic execution into a sequence of information flows, each annotated with the five CI parameters, and traces violations to their point of origin. We present AgentSCOPE, a benchmark of 62 multi-tool scenarios across eight regulatory domains with ground truth at every pipeline stage. Our evaluation across seven state-of-the-art LLMs show that privacy violations in the pipeline occur in over 80% of scenarios, even when final outputs appear clean (24%), with most violations arising at the tool-response stage where APIs return sensitive data indiscriminately. These results indicate that output-level evaluation alone substantially underestimates the privacy risk of agentic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。