arXiv:2503.01630cs.LGcs.AI2025-03被引 6

大模型可能存储个人数据,需承担法律保护责任。

Machine Learners Should Acknowledge the Legal Implications of Large Language Models as Personal Data

  • 模型推理时会泄露训练中的个人数据
  • 部分模型本身可被视为个人数据
  • 研究者应全程关注法律合规问题

大型语言模型(LLMs)在训练过程中会不同程度地记忆数据。即使仅记忆少量个人数据,一旦可识别特定个体,即属于数据保护法范畴。即便训练完成,模型仍受欧盟《通用数据保护条例》(GDPR)约束。本文指出:(1)模型在推理阶段可能输出训练数据,无论是原文还是泛化形式;(2)某些模型因包含可识别信息,可被认定为个人数据,从而触发访问、更正、删除等数据主体权利;(3)机器学习研究人员必须在数据收集、模型训练到公开发布(如GitHub或Hugging Face)的全周期中考虑法律影响;(4)提出应对策略,推动法律与技术的协同。研究强调需加强法律界与机器学习社区的对话。

原文摘要 · Abstract (English)

Does GPT know you? The answer depends on your level of public recognition; however, if your information was available on a website, the answer could be yes. Most Large Language Models (LLMs) memorize training data to some extent. Thus, even when an LLM memorizes only a small amount of personal data, it typically falls within the scope of data protection laws. If a person is identified or identifiable, the implications are far-reaching. The LLM is subject to EU General Data Protection Regulation requirements even after the training phase is concluded. To back our arguments: (1.) We reiterate that LLMs output training data at inference time, be it verbatim or in generalized form. (2.) We show that some LLMs can thus be considered personal data on their own. This triggers a cascade of data protection implications such as data subject rights, including rights to access, rectification, or erasure. These rights extend to the information embedded within the AI model. (3.) This paper argues that machine learning researchers must acknowledge the legal implications of LLMs as personal data throughout the full ML development lifecycle, from data collection and curation to model provision on e.g., GitHub or Hugging Face. (4.) We propose different ways for the ML research community to deal with these legal implications. Our paper serves as a starting point for improving the alignment between data protection law and the technical capabilities of LLMs. Our findings underscore the need for more interaction between the legal domain and the ML community.

大模型隐私保护法律合规

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。