arXiv:2507.03034cs.LGcs.AI2025-07被引 29

为生成式AI时代的数据保护构建四层框架,明确防护边界。

Rethinking Data Protection in the (Generative) Artificial Intelligence Era

  • 提出非可用性、隐私保护、可追溯性、可删除性四层保护框架
  • 揭示现有监管在模型权重和生成内容上的保护盲区
  • 帮助开发者与监管者平衡数据效用与控制权

生成式人工智能时代深刻改变了数据的内涵与价值。数据不再仅是静态内容,而是贯穿模型训练、提示输入与输出部署全过程的关键要素。传统数据保护理念已显不足,且需保护范围界定不清。若未妥善保护AI系统中的数据,将对社会和个人造成严重损害。为此,本文提出一个包含非可用性、隐私保护、可追溯性和可删除性的四层分类体系,涵盖现代生成式AI模型与系统中多样化的保护需求。该框架系统呈现了数据效用与控制权之间的权衡关系,覆盖从训练数据集、模型参数、系统提示到生成内容的全生命周期。我们分析了各层级代表性技术方案,并指出当前监管中存在的盲点,导致关键资产暴露。本框架为未来人工智能技术与治理提供结构化视角,推动可信数据实践,强调重新审视生成式AI时代数据保护的紧迫性,并为开发者、研究者与监管机构提供及时指导。

原文摘要 · Abstract (English)

The (generative) artificial intelligence (AI) era has profoundly reshaped the meaning and value of data. No longer confined to static content, data now permeates every stage of the AI lifecycle from the training samples that shape model parameters to the prompts and outputs that drive real-world model deployment. This shift renders traditional notions of data protection insufficient, while the boundaries of what needs safeguarding remain poorly defined. Failing to safeguard data in AI systems can inflict societal and individual, underscoring the urgent need to clearly delineate the scope of and rigorously enforce data protection. In this perspective, we propose a four-level taxonomy, including non-usability, privacy preservation, traceability, and deletability, that captures the diverse protection needs arising in modern (generative) AI models and systems. Our framework offers a structured understanding of the trade-offs between data utility and control, spanning the entire AI pipeline, including training datasets, model weights, system prompts, and AI-generated content. We analyze representative technical approaches at each level and reveal regulatory blind spots that leave critical assets exposed. By offering a structured lens to align future AI technologies and governance with trustworthy data practices, we underscore the urgency of rethinking data protection for modern AI techniques and provide timely guidance for developers, researchers, and regulators alike.

数据保护生成式AI治理框架隐私

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。