arXiv:2607.05479cs.CRcs.AI2026-07

解析生成式AI三类数据处理模式的保密风险

Privilege and confidentiality in generative AI workflows

  • 区分模型参数、会话上下文、知识库三种数据存储方式
  • 指出每种模式均存在隐蔽的机密泄露风险
  • 为法律从业者提供合规治理新标准参考

生成式AI系统通过三种方式处理客户数据:训练与记忆存于模型参数中,实时会话中暂存于上下文窗口,以及通过检索增强生成(RAG)存于知识数据库。每种模式均带来不同且常反直觉的保密与法律职业特权风险,需对应特定治理措施。基于首两起英美判例(UK and Munir v Secretary of State for the Home Department 与 United States v Heppner),结合传统特权法理与近期计算机科学研究成果,本文以实务者可理解的方式解释这三种数据处理模式,并分析其法律后果。进而将分析置于英格兰和威尔士律师监管框架及普通专业过失原则下,论证有效信息治理标准正在改变。虽主要面向SRA监管下的执业者,但本数据治理分析可扩展至任何依赖可证明保密性保护特权或职业秘密的司法管辖区。最终目标是帮助法律服务从业者识别生成式AI中的关键数据泄露风险,推动对客户数据及敏感材料更负责任的部署。

原文摘要 · Abstract (English)

Generative AI (GenAI) systems store and process client data in three distinct ways: in the model's parameters through training and memorisation, in the context window during a live session, and in knowledge databases for retrieval-augmented generation (RAG). Each mode creates different and often counter-intuitive risks to confidentiality and legal professional privilege, and each calls for specific governance responses. Drawing on the first English and American decisions to address privilege and generative AI, UK and Munir v Secretary of State for the Home Department and United States v Heppner, on the orthodox privilege authorities against which those decisions must be read, and on recent computer science research, we explain the three modes of data storage and processing in terms accessible to practitioners and analyse the legal consequences of each. We then situate the analysis within the regulatory framework governing solicitors in England and Wales and within the ordinary principles of professional negligence, arguing that the standard of effective information governance (and with it the benchmark against which negligence and misconduct will be measured) is changing. Although we write primarily for SRA-regulated practitioners, our data-governance analysis is framed to extend to any jurisdiction in which the protection of privilege or professional secrecy depends on demonstrable confidentiality. The ultimate aim of this article is to help legal services professionals understand salient data leakage risks in GenAI systems and thereby facilitate a more responsible deployment of GenAI on client data and other sensitive material.

生成式AI数据隐私法律科技

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。