arXiv:2604.21308cs.CRcs.CL2026-04ACL被引 7

测试企业大模型在工作流中的信息泄露风险,发现性能越强越容易泄密。

CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents

论文配图:CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents
图 1 · 摘自论文原文
  • 构建跨五类信息流的基准测试,评估模型是否能准确传递内容又隐藏敏感信息。
  • 前沿模型隐私泄漏率高达15.8%至50.9%,最高泄露达26.7%。
  • 适合关注企业级AI安全、数据合规与上下文管理的研究者和工程师。

企业大模型可显著提升工作效率,但其依赖内部上下文执行任务的能力也带来了敏感信息泄露的新风险。我们提出CI-Work,一个基于上下文完整性(CI)的基准,模拟企业工作流中五个信息流向,在密集检索场景下评估模型能否传递必要内容同时屏蔽敏感上下文。对前沿模型的评估显示,隐私失败普遍存在(违规率15.8%–50.9%,泄露最高达26.7%),并揭示一个反直觉的权衡:任务效用越高,隐私违规越严重。此外,企业数据规模庞大及用户行为复杂性进一步加剧此风险。单纯增大模型规模或深化推理无法解决问题。结论表明,保障企业工作流需从模型中心转向上下文中心的架构范式。

原文摘要 · Abstract (English)

Enterprise LLM agents can dramatically improve workplace productivity, but their core capability, retrieving and using internal context to act on a user's behalf, also creates new risks for sensitive information leakage. We introduce CI-Work, a Contextual Integrity (CI)-grounded benchmark that simulates enterprise workflows across five information-flow directions and evaluates whether agents can convey essential content while withholding sensitive context in dense retrieval settings. Our evaluation of frontier models reveals that privacy failures are prevalent (violation rates range from 15.8%-50.9%, with leakage reaching up to 26.7%) and uncovers a counterintuitive trade-off critical for industrial deployment: higher task utility often correlates with increased privacy violations. Moreover, the massive scale of enterprise data and potential user behavior further amplify this vulnerability. Simply increasing model size or reasoning depth fails to address the problem. We conclude that safeguarding enterprise workflows requires a paradigm shift, moving beyond model-centric scaling toward context-centric architectures.

企业AI隐私安全上下文管理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。